Concurrency Basics
Topic 3 of 7, with 3 concept checks. Threading vs multiprocessing vs asyncio, and the GIL
Choose the concurrency model that matches the wait
Concurrent work
Separate CPU work from blocking waits, shared-memory threads from isolated processes, and concurrency from parallelism before selecting a Python execution model.
Core lesson 01
The Global Interpreter Lock (GIL) allows only one thread to execute Python bytecode at a time in CPython, so threading gives no speedup for CPU-bound work — but still helps for I/O-bound work.
CPython's memory management isn't thread-safe by default, so the interpreter uses a single global lock to protect it — only the thread holding the GIL can run Python bytecode. Threads still get interleaved (the GIL is released periodically), which is why threading works great for I/O-bound tasks (a thread waiting on a network call releases the GIL so another can run) but doesn't parallelize CPU-bound number crunching across cores the way a purely thread-based language would.
import threading, time
def cpu_bound():
total = 0
for i in range(20_000_000):
total += i
return total
start = time.time()
threads = [threading.Thread(target=cpu_bound) for _ in range(2)]
[t.start() for t in threads]
[t.join() for t in threads]
print("threaded:", time.time() - start)
# roughly the same or worse than running cpu_bound() twice sequentially,
# because the GIL prevents true parallel execution of Python bytecodeWhat to remember
What is the GIL, and why does it mean threading doesn't speed up CPU-bound Python code?
Common footguns
- Using threading to 'speed up' a CPU-bound computation and being confused why it doesn't get faster — use multiprocessing instead, which sidesteps the GIL with separate processes.
Core lesson 02
multiprocessing runs separate processes, each with its own GIL, giving true CPU parallelism at the cost of higher memory use and needing to serialize (pickle) data passed between processes.
Since each process in multiprocessing has its own interpreter and GIL, CPU-bound work genuinely runs in parallel across cores — ideal for number crunching or image processing. The cost: processes don't share memory, so passing data between them requires pickling it, which adds overhead and means not all objects can even be shared. Threading, by contrast, shares memory directly, which is lighter weight and better suited to I/O-bound work, where the GIL is released during the wait anyway.
from multiprocessing import Pool
def square(n):
return n * n
if __name__ == "__main__":
with Pool(processes=4) as pool:
results = pool.map(square, range(10))
print(results) # computed across 4 real processesWhat to remember
When would you choose multiprocessing over threading, and what's the tradeoff?
Common footguns
- Forgetting the if __name__ == '__main__': guard when using multiprocessing — without it on some platforms, child processes re-import and re-run the whole script.
Core lesson 03
asyncio runs everything on a single thread using an event loop; coroutines cooperatively yield at each await, so many I/O-bound tasks can be 'in flight' concurrently without thread/process overhead — but one blocking call freezes everything.
Where threading relies on OS preemption and multiprocessing uses separate OS processes, asyncio achieves concurrency within a single thread: an event loop runs one coroutine until it hits an await on something not yet ready, then switches to another ready coroutine. This is extremely lightweight and ideal for I/O-bound workloads with many concurrent waits — but any single blocking, non-async call stalls the entire event loop, since there's no OS-level preemption to save you.
import asyncio
async def fetch(name, delay):
print(f"{name} starting")
await asyncio.sleep(delay)
print(f"{name} done")
return name
async def main():
results = await asyncio.gather(fetch("A", 2), fetch("B", 1))
print(results)
asyncio.run(main())
# both start immediately; total time is ~2s, not 3s -- they ran concurrentlyWhat to remember
How does asyncio differ from both threading and multiprocessing?
Common footguns
- Calling a blocking function like time.sleep() (instead of await asyncio.sleep()) inside a coroutine — this freezes the entire event loop.
Threading: [T1-run][T2-run][T1-run][T2-run]... (GIL: 1 at a time, OS-preempted)
Multiprocessing:[P1-run||||||||||] (separate GILs, true parallel)
[P2-run||||||||||]
Asyncio: [coro A][await][coro B][await][coro A resumes]... (1 thread, cooperative)Python lab
Browser Python lab
Runtime · idle
Python loads on your first run. Your code stays in this browser.
Best practices
- Use asyncio or threading for I/O-bound work; use multiprocessing for CPU-bound work.
- Never mix blocking calls into async code without wrapping them (e.g. loop.run_in_executor) — one blocking call stalls the whole event loop.
- Guard multiprocessing entry points with if __name__ == '__main__': for cross-platform safety.
- Prefer asyncio.gather or TaskGroups over manually awaiting many coroutines sequentially, which loses the concurrency benefit.
Apply the concept in Interview practice
Print in OrdereasyLeetCode #1114 · O(1) synchronization overhead
Use two Semaphores (or Events): release the one for the next method at the end of the current one, and have each method acquire its semaphore before running.
Open problemPrint FooBar AlternatelymediumLeetCode #1115 · O(n) synchronized steps
Use two semaphores initialized so foo starts first; each method waits on its own semaphore and signals the other's after printing.
Open problemThe Dining PhilosophersmediumLeetCode #1226 · O(1) per philosopher per round
Have each philosopher pick up the lower-numbered fork first (a fixed global ordering) to prevent circular wait — the classic deadlock-avoidance trick.
Open problemConcept checks
What is the GIL, and why does it mean threading doesn't speed up CPU-bound Python code?
Hint
Only one thread can execute Python bytecode at a time, no matter how many CPU cores you have.
It's a lock around the interpreter itself, not your code specifically.
Answer
The Global Interpreter Lock (GIL) allows only one thread to execute Python bytecode at a time in CPython, so threading gives no speedup for CPU-bound work — but still helps for I/O-bound work.
CPython's memory management isn't thread-safe by default, so the interpreter uses a single global lock to protect it — only the thread holding the GIL can run Python bytecode. Threads still get interleaved (the GIL is released periodically), which is why threading works great for I/O-bound tasks (a thread waiting on a network call releases the GIL so another can run) but doesn't parallelize CPU-bound number crunching across cores the way a purely thread-based language would.
import threading, time
def cpu_bound():
total = 0
for i in range(20_000_000):
total += i
return total
start = time.time()
threads = [threading.Thread(target=cpu_bound) for _ in range(2)]
[t.start() for t in threads]
[t.join() for t in threads]
print("threaded:", time.time() - start)
# roughly the same or worse than running cpu_bound() twice sequentially,
# because the GIL prevents true parallel execution of Python bytecodeWatch out
- Using threading to 'speed up' a CPU-bound computation and being confused why it doesn't get faster — use multiprocessing instead, which sidesteps the GIL with separate processes.
When would you choose multiprocessing over threading, and what's the tradeoff?
Hint
One gets true parallelism across CPU cores; the other has much cheaper communication.
Processes don't share memory by default — that's the cost of escaping the GIL.
Answer
multiprocessing runs separate processes, each with its own GIL, giving true CPU parallelism at the cost of higher memory use and needing to serialize (pickle) data passed between processes.
Since each process in multiprocessing has its own interpreter and GIL, CPU-bound work genuinely runs in parallel across cores — ideal for number crunching or image processing. The cost: processes don't share memory, so passing data between them requires pickling it, which adds overhead and means not all objects can even be shared. Threading, by contrast, shares memory directly, which is lighter weight and better suited to I/O-bound work, where the GIL is released during the wait anyway.
from multiprocessing import Pool
def square(n):
return n * n
if __name__ == "__main__":
with Pool(processes=4) as pool:
results = pool.map(square, range(10))
print(results) # computed across 4 real processesWatch out
- Forgetting the if __name__ == '__main__': guard when using multiprocessing — without it on some platforms, child processes re-import and re-run the whole script.
How does asyncio differ from both threading and multiprocessing?
Hint
It's still single-threaded — concurrency comes from cooperative switching, not parallelism.
Tasks voluntarily yield control at await points instead of being preempted.
Answer
asyncio runs everything on a single thread using an event loop; coroutines cooperatively yield at each await, so many I/O-bound tasks can be 'in flight' concurrently without thread/process overhead — but one blocking call freezes everything.
Where threading relies on OS preemption and multiprocessing uses separate OS processes, asyncio achieves concurrency within a single thread: an event loop runs one coroutine until it hits an await on something not yet ready, then switches to another ready coroutine. This is extremely lightweight and ideal for I/O-bound workloads with many concurrent waits — but any single blocking, non-async call stalls the entire event loop, since there's no OS-level preemption to save you.
import asyncio
async def fetch(name, delay):
print(f"{name} starting")
await asyncio.sleep(delay)
print(f"{name} done")
return name
async def main():
results = await asyncio.gather(fetch("A", 2), fetch("B", 1))
print(results)
asyncio.run(main())
# both start immediately; total time is ~2s, not 3s -- they ran concurrentlyThreading: [T1-run][T2-run][T1-run][T2-run]... (GIL: 1 at a time, OS-preempted)
Multiprocessing:[P1-run||||||||||] (separate GILs, true parallel)
[P2-run||||||||||]
Asyncio: [coro A][await][coro B][await][coro A resumes]... (1 thread, cooperative)Watch out
- Calling a blocking function like time.sleep() (instead of await asyncio.sleep()) inside a coroutine — this freezes the entire event loop.