OS MULTITHREADING

Let’s assume you wrote your first C++ program. Now the next step is to use a compiler to compile it into an executable (i.e. an .exe file), which is the binary that your computer understands. Then you run that executable to perform the intended task in the C++ program.
Here, the execution of that executable is called a Process.
So we can say that running a program on a computer creates a process.
Now a Thread is like a lightweight process.
Let’s take a case where:
You want a program that takes an image as input and then uploads it to 2 different cloud storages.
If you write the entire program using a single thread, the CPU will wait for the image to be uploaded to the first cloud storage and then continue with the second cloud storage. In between, the CPU may have very little useful work to do because the program is waiting for the network response. This wastes time.
To overcome this, we can move to a Multithreading architecture.
Here, in the program, we segregate the 2 upload tasks into 2 separate threads, allowing both uploads to make progress independently. This can reduce the overall execution time, especially because uploading an image is an I/O-bound task.
To understand this further, you need to know these concepts:
Process Vs Thread
Processes are isolated from each other. That is, they have their own private virtual memory and address space.
Whereas threads within the same process share the process’s memory and address space.
This is one of the main reasons why threads are considered lightweight compared to processes.
Concurrency Vs Parallelism
Concurrency is about multiple tasks making progress during overlapping periods of time. On a single CPU core, this can be achieved through context switching, where the OS switches between threads.
Parallelism means multiple tasks are actually executing at the same time, which is possible when multiple CPU cores execute different threads simultaneously.
For example:
Think of Concurrency like a chef working in the kitchen. He puts some potatoes in the fryer and, in the meantime, chops some veggies to serve alongside them.
He is handling multiple tasks by switching between them while one task is waiting.
In Parallelism, imagine 2 chefs working on the tasks separately at the same time.
Now let’s come back to the image upload case study.
On a Single-Core System (1 CPU Core)
The CPU relies on context switching.
It starts sending image packets to Cloud 1 using Thread A. While Thread A is waiting for Cloud 1 to acknowledge receipt, the OS can pause Thread A and switch to Thread B, which starts sending packets to Cloud 2.
The OS keeps switching between them so rapidly that both uploads appear to be progressing at the same time.
On a Multi-Core System (more than 1 CPU Core)
The OS sees two runnable threads and can schedule them on different CPU cores.
Core 1 takes Thread A (Cloud 1) and Core 2 takes Thread B (Cloud 2).
Both CPU cores can execute instructions for the two threads simultaneously.
This is true parallelism.
# Note:
The above case is an I/O-bound task.
Most of the time, the threads are waiting for the network and the cloud servers to respond. So adding another thread doesn’t necessarily make the CPU work twice as fast.
In fact, the difference between a single-core and multi-core system may not be very significant for this particular workload. The actual performance depends heavily on the network, cloud servers, latency, bandwidth, and how the program handles I/O.
Now let’s look at CPU-bound tasks, such as:
- Rendering 3D graphics
- Training ML models
- Encoding videos
Here, the hardware is doing heavy computation.
1-Core System :
Multithreading a CPU-bound task on a single core does not provide true parallelism.
The CPU has to keep switching between threads, and those context switches introduce overhead.
So, depending on the workload, multithreading can provide little benefit or even make the program slower.
Multi-Core System:
Now we can actually take advantage of multiple CPU cores.
For example, if we have 2 CPU cores and 2 independent CPU-bound threads, Core 1 can execute Thread A while Core 2 executes Thread B.
Both threads can execute simultaneously.
So, ideally, the execution time can be close to half compared to running the same work sequentially on one core.
However, the actual speedup depends on things like synchronization, memory bandwidth, workload distribution, cache behavior, and the number of available cores.