What Is the GIL? Python Threads vs Processes vs asyncio
The GIL lets one thread run Python bytecode at a time. Why threads help I/O-bound but not CPU-bound code, and when to use threads, processes or asyncio.
- Course: Python study plan
- Module: Concurrency — threads, processes and asyncio
- Kind: Lesson
- Reading time: 14 min
- Runtime: CPython 3.11
What is the GIL in Python?
The global interpreter lock (GIL) is the mutex in CPython that lets only one thread execute Python bytecode at a time, protecting the reference counts that manage memory. A thread releases it while blocked on I/O, and some C extensions such as NumPy release it during heavy loops. Threads therefore speed up code that waits, but not CPU-bound pure-Python code.
Lesson
Python has three ways to do more than one thing at a time, and choosing among them starts from one fact about CPython: the global interpreter lock allows only one thread to execute Python bytecode at any moment. Threads therefore do not speed up CPU-bound Python code, but they do speed up code that waits — on the network, the disk, a subprocess — because a waiting thread releases the lock. Processes sidestep the lock with separate interpreters; asyncio gives waiting a single-threaded, explicit form. This lesson explains the GIL and its consequences, classifies work as CPU-bound or I/O-bound, gives the decision table for the three models, and states the rule that keeps concurrent programs testable and judgeable: collect results in a defined order, then print.
The GIL
CPython's memory management is reference counting, and updating a count from two threads at once would corrupt it; the GIL is the mutex that prevents that by serialising bytecode execution. A thread holds the lock while running Python code, and releases it when it blocks on I/O (a socket read, a file read, time.sleep) or when a C extension explicitly drops it (NumPy's numerical loops, hashlib, zlib). Every 5 ms (sys.getswitchinterval()) a running thread is asked to yield so others can run.
Consequences:
- Two threads doing pure-Python arithmetic take as long as one thread doing both — sometimes longer, from contention.
- Two threads each waiting on a network response overlap almost perfectly: while one waits, the other runs.
- Thread switches can happen between any two bytecodes, so
count += 1from two threads is a race — the read and the write are separate operations (lesson 2).
Python 3.12 gave each sub-interpreter its own GIL and 3.13 ships an experimental free-threaded build (--disable-gil); both are reading for this track — on 3.11 the GIL is simply there.
CPU-bound versus I/O-bound
| Work | Examples | Bottleneck | Model |
|---|---|---|---|
| CPU-bound | parsing, hashing in Python, image processing in pure Python, simulation | the processor | processes (multiprocessing, ProcessPoolExecutor), or a C extension that releases the GIL |
| I/O-bound, few tasks | reading a handful of files, a few HTTP calls, a subprocess | waiting | threads (threading, ThreadPoolExecutor) |
| I/O-bound, many tasks | thousands of connections, a chat server, a scraper with high concurrency | waiting, at scale | asyncio |
| Mixed | fetch then compute | both | asyncio or threads for the fetch, a process pool for the compute |
The test is: if the program had infinitely fast I/O, would it still be slow? Then it is CPU-bound and threads will not help. time.perf_counter around the work versus around the waiting tells you which.
The three models
Threads share memory: one process, several threads, every object visible to all. Cheap to start, simple to reason about for a handful of workers, and dangerous exactly because everything is shared — the lock discipline of the next lesson is mandatory. Blocking calls just work.
Processes share nothing: each is a separate interpreter with its own GIL and memory. True parallelism for CPU work at the cost of start-up time (tens of milliseconds each), and of passing data by pickling it — arguments and results must be picklable, and copying a large object to a worker can cost more than the work. concurrent.futures.ProcessPoolExecutor hides most of the mechanics.
asyncio is one thread with an event loop: coroutines (async def) run until they await something that is not ready, hand control back to the loop, and are resumed when it is. No parallelism at all — one thing runs at a time — but tens of thousands of concurrent waits with almost no memory, and switches happen only at await, which makes shared state safe between awaits. The cost: every blocking call in the program must be an awaitable one; a single time.sleep(1) freezes the whole loop.
The determinism rule
Concurrent programs are nondeterministic in scheduling: which thread runs first, which task finishes first, the order in which messages interleave. Output that depends on that order differs from run to run and cannot be tested or judged. The rule for every exercise in this module, and for well-designed programs generally:
- Give each worker a slot for its result — an index into a list, a key in a dict, a return value collected by
mapin submission order. - Join every worker (or
gatherevery task) before reading results. - Print the aggregated, ordered results after the join, from the main thread.
executor.map returns results in the order the inputs were submitted regardless of completion order; asyncio.gather does the same for coroutines; as_completed does not, and its results must be sorted before printing. A worker that prints directly interleaves with others and the interleaving is arbitrary.
What to say in an interview
The GIL serialises bytecode, so threads are for I/O and processes for CPU; asyncio is cooperative single-threaded concurrency for many I/O waits; a race is two threads touching one object between bytecodes; a lock makes a compound operation atomic; a deadlock is two locks acquired in different orders; the fix for shared state is to not share it — queues, immutable messages, results returned rather than written. Those six sentences are the whole of what most interviewers want, and the rest of this module is the practice behind them.
Pitfalls
- Threads for a CPU-bound loop, then surprise that nothing got faster.
- A blocking call (
requests.get,time.sleep, a file read) inside anasync def. - Printing from workers and expecting a stable order.
- Reading results before the join.
- Passing an unpicklable object (a lambda, an open file, a lock) to a process pool.
- Assuming
x += 1is atomic.
Key takeaways
- The GIL lets one thread run Python bytecode at a time; it is released during I/O and in some C extensions.
- CPU-bound work wants processes (or a GIL-releasing extension); I/O-bound work wants threads for a few tasks and asyncio for many.
- Threads share memory and need locks; processes share nothing and pickle data; asyncio switches only at
awaitand must never block. - Collect results into ordered slots, join everything, then print — the only way concurrent output is deterministic.
x += 1is not atomic; a race is a switch between its read and its write.
Common questions
Why don't Python threads speed up CPU-bound code?
Only the thread holding the GIL runs bytecode, so two threads doing pure-Python arithmetic take as long as one thread doing both, sometimes longer from contention. CPU-bound work needs processes, through multiprocessing or ProcessPoolExecutor, or a C extension that releases the GIL.
When should I use threads, processes or asyncio in Python?
Use processes for CPU-bound work, threads for I/O-bound work with a handful of tasks and blocking libraries, and asyncio for many concurrent waits such as thousands of connections. The test: if the I/O were instantly fast and the program were still slow, it is CPU-bound.
Is x += 1 atomic in Python?
No. It reads the value, adds one and writes it back as separate bytecode steps, and a thread switch can happen between them, so two threads can read the same value and one increment is lost. Protect the compound operation with a Lock.
Is Python getting rid of the GIL?
Python 3.12 gave each sub-interpreter its own GIL, and Python 3.13 added an experimental free-threaded build that runs without it. The default CPython build still has the GIL, and on Python 3.11 and earlier there is no way to turn it off.
How do you make concurrent Python output deterministic?
Give each worker its own result slot, join every thread or gather every task before reading, then print the ordered results from the main thread. executor.map and asyncio.gather return results in submission order; as_completed does not, so sort what it yields.
Exercises
Classify the work
Read lines <name> <cpu_ms> <io_ms> <tasks> until EOF. Let the I/O share be io_ms / (cpu_ms + io_ms). A share of at least 0.75 is io-bound and recommends threads when tasks is at most 50, otherwise asyncio; a share of at most 0.25 is cpu-bound and recommends processes; anything else is mixed and recommends threads+processes. Print <name>: <class> -> <model> per line, then cpu-bound <a>, io-bound <b>, mixed <c>.
Input: one job per line. Output: one line per job, then the summary.
parse 400 0 8
crawl 10 190 500
fetch 20 180 6
render 100 100 4
prints
parse: cpu-bound -> processes
crawl: io-bound -> asyncio
fetch: io-bound -> threads
render: mixed -> threads+processes
cpu-bound 1, io-bound 2, mixed 1Slots after join
Read one line of integers. Start one thread per integer, named worker-<i>, whose target computes sum(k*k for k in range(1, n + 1)) and stores (threading.current_thread().name, value) in results[i] — its own slot. Join every thread, then print <name> n=<n> -> <value> in slot order and total <sum of values>. Nothing may be printed from inside a worker.
Input: one line of integers. Output: one line per slot, then the total.
3 5 10
prints
worker-0 n=3 -> 14
worker-1 n=5 -> 55
worker-2 n=10 -> 385
total 454In this module: Concurrency — threads, processes and asyncio
- The GIL and the three models — threads, processes, asyncio (this lesson)
- Threads — Thread, Lock, Event, Queue and the race you must see once
- concurrent.futures — executors, futures, map and as_completed
- multiprocessing — Process, Pool, pickling, queues and shared state
- asyncio basics — coroutines, await, tasks and gather
- asyncio patterns — TaskGroup, queues, semaphores, async iteration and bridging
- Checkpoint — Concurrency
← Checkpoint — Decorators, descriptors and the data model · Threads — Thread, Lock, Event, Queue and the race you must see once →