Python Multiprocessing Explained: Pool, Pickling and Queues
Python multiprocessing runs CPU-bound work in separate processes, each with its own GIL. Process and Pool, chunksize, pickling, the main guard, shared state.
- Course: Python study plan
- Module: Concurrency — threads, processes and asyncio
- Kind: Lesson
- Reading time: 13 min
- Runtime: CPython 3.11
What is multiprocessing in Python?
Multiprocessing in Python runs functions in separate interpreter processes, each with its own memory and its own GIL, which gives true parallelism for CPU-bound work. The multiprocessing module provides Process and Pool; everything that crosses to a worker, meaning the function, its arguments and its result, is pickled, and starting a process costs tens of milliseconds.
Lesson
multiprocessing is the lower-level module under ProcessPoolExecutor: it starts interpreter processes, runs functions in them, and moves data between them. It exists because processes are how CPU-bound Python gets parallel, and its constraints — everything crosses the boundary by pickling, the main module is re-imported by workers, start-up costs tens of milliseconds — are the constraints of that design. This lesson covers Process and Pool, the map/imap/starmap family, what can and cannot be pickled, the start methods and the main guard, Queue/Pipe for messages, Value/Array/Manager for shared state in outline, and the cost model that decides when a pool pays.
Process
from multiprocessing import Process, Queue
def work(n, out):
out.put((n, sum(i * i for i in range(n))))
if __name__ == "__main__":
out = Queue()
procs = [Process(target=work, args=(n, out)) for n in (10**5, 2 * 10**5)]
for p in procs: p.start()
results = sorted(out.get() for _ in procs) # collect, then order by the key
for p in procs: p.join()
print(results)
Same shape as Thread — target, args, start, join — but the function runs in a new interpreter with its own memory. Process has pid, exitcode, is_alive(), terminate(). Results come back through a Queue (or a return value via a pool), never through a shared variable: a global the child modifies is the child's copy.
Pool
from multiprocessing import Pool
def square(n):
return n * n
def power(base, exp):
return base ** exp
if __name__ == "__main__":
with Pool(processes=4) as pool:
pool.map(square, range(10)) # ordered list; blocks
pool.map(square, range(10**6), chunksize=1000) # batch small tasks
list(pool.imap(square, range(10))) # lazy, ordered
list(pool.imap_unordered(square, range(10))) # lazy, completion order — sort it
pool.starmap(power, [(2, 3), (3, 2)]) # unpack argument tuples
pool.apply_async(square, (7,)).get(timeout=5) # one job, a result handle
Pool.map is ProcessPoolExecutor.map with more knobs; imap streams results; starmap takes argument tuples. chunksize is the lever that matters: each task costs a pickle round trip and a queue hop, so a million tiny tasks need batching into thousands. The with block terminates the workers on exit; pool.close(); pool.join() is the explicit form that waits for them.
Pickling
Everything that crosses to a worker — the function, its arguments, its result — is serialised with pickle. That works for built-in types, dataclasses, plain classes defined at module level, and functions defined at module level; it fails for lambdas, nested functions, bound methods of unpicklable objects, open files, locks, database connections and generators. The error is PicklingError or AttributeError: Can't pickle local object. The fixes: define the worker function at module level; pass data, not resources (open the file inside the worker); use functools.partial on a module-level function instead of a lambda; and for large arrays, prefer shared memory (multiprocessing.shared_memory, NumPy) or have the worker load its own data rather than shipping it.
Start methods and the main guard
A worker process must obtain the code to run. With fork (Linux default) it inherits a copy of the parent's memory; with spawn (macOS and Windows default) it starts a fresh interpreter and imports the main module to find the function. That import re-executes top-level code, so a script that creates a pool at top level spawns workers that create pools that spawn workers. Hence the rule: everything that starts processes lives under if __name__ == "__main__":. multiprocessing.set_start_method("spawn") chooses explicitly; fork is faster but unsafe with threads already running.
Queue and Pipe
multiprocessing.Queue is the cross-process version of queue.Queue: put/get with blocking and timeouts, pickling underneath. Pipe() returns two connected Connection objects for two-party messaging (send/recv). Both are for streaming work and results between long-running processes; a pool's map is simpler when the work is a batch.
Shared state, in outline
Value("i", 0) and Array("d", 10) allocate typed memory shared between processes, with a .get_lock() for compound updates; Manager() starts a server process holding proxies for a dict or list that several processes may modify (slow, but general). Both are last resorts: the design that scales is no shared state — each worker computes from its inputs and returns a result, and the parent aggregates.
The cost model
Starting a process costs 10–100 ms; pickling costs time proportional to the data's size; a result must be pickled back. A pool pays when each task is at least milliseconds of CPU work and its data is small relative to the work, or when the data is loaded inside the worker. It does not pay for a hundred tiny tasks, for tasks that ship large arguments, or for I/O-bound work (threads or asyncio are cheaper). On the judge two cores exist, so a pool of two can genuinely halve a CPU-bound computation — and every result must still be collected and ordered before printing.
Pitfalls
- No main guard: recursive spawning on macOS/Windows.
- A lambda or nested function as the worker.
- Expecting a global modified in a worker to change in the parent.
- Shipping a large object to every task instead of loading it in the worker.
- A million tiny tasks without
chunksize. imap_unorderedoutput printed in arrival order.
Key takeaways
Process(target=…)mirrorsThreadbut in a separate interpreter; results return through aQueueor a pool, never shared variables.Pool.map/imap/starmaprun batches;chunksizeamortises the per-task pickling cost;imap_unorderedmust be sorted.- Workers and their arguments are pickled: module-level functions, plain data, no lambdas or resources.
- Workers re-import the main module under
spawn, so process creation lives under the main guard. - Shared memory and managers exist; returning results and aggregating in the parent is the design to prefer.
Common questions
What is the difference between multiprocessing and multithreading in Python?
Threads share one process's memory and one GIL, so they overlap waiting but not Python computation. Processes share nothing: each has its own interpreter and GIL, giving real parallelism for CPU work, at the cost of start-up time and pickling every argument and result.
Why do I get "Can't pickle local object" in multiprocessing?
The worker function or one of its arguments cannot be pickled: lambdas, nested functions, open files, locks, connections and generators cannot. Define the worker at module level, pass plain data rather than resources, and use functools.partial on a module-level function instead of a lambda.
What does chunksize do in Pool.map?
It batches tasks so that each pickle round trip and queue hop carries many items instead of one. Every task has a fixed overhead, so a million tiny tasks without batching can be slower than a plain loop; a chunksize in the thousands amortises it.
Why doesn't a global changed in a worker process update the parent?
Each worker has its own copy of memory, so a global it modifies is the child's copy. Return results through Pool.map, a Queue or a Pipe and aggregate them in the parent; Value, Array and Manager share state, but only as a last resort.
What is the difference between fork and spawn in multiprocessing?
With fork, the Linux default up to Python 3.13, a worker inherits a copy of the parent's memory. With spawn, the default on macOS and Windows, it starts a fresh interpreter and imports the main module, which is why code that starts processes must sit under if __name__ == '__main__':.
Exercises
starmap and the unordered check
Read lines base exp mod until EOF. Write a module-level pow_mod(base, exp, mod) returning pow(base, exp, mod) and a module-level indexed(job) that takes (i, base, exp, mod) and returns (i, pow_mod(base, exp, mod)). Under the main guard, open Pool(2), compute the results with starmap and print <base>^<exp> mod <mod> = <r> per job in input order; then run the indexed jobs through imap_unordered, sort the pairs by index, and print unordered matches <True/False> comparing them with the ordered results.
Input: one job per line. Output: one line per job, then the comparison.
2 10 1000
3 4 5
7 0 13
prints
2^10 mod 1000 = 24
3^4 mod 5 = 1
7^0 mod 13 = 1
unordered matches TrueTwo processes and a queue
Read one line of integers and split it into two parts: the first (n + 1) // 2 numbers and the rest. Start one multiprocessing.Process per part running a module-level part_sum(index, chunk, q) that puts (index, sum(chunk)) on a multiprocessing.Queue. In the parent, collect both results from the queue, join the processes, sort by index, and print part <i> sum <s> for each and total <sum>.
Input: one line of integers. Output: three lines.
1 2 3 4 5
prints
part 0 sum 6
part 1 sum 9
total 15In this module: Concurrency — threads, processes and asyncio
- The GIL and the three models — threads, processes, asyncio
- Threads — Thread, Lock, Event, Queue and the race you must see once
- concurrent.futures — executors, futures, map and as_completed
- multiprocessing — Process, Pool, pickling, queues and shared state (this lesson)
- asyncio basics — coroutines, await, tasks and gather
- asyncio patterns — TaskGroup, queues, semaphores, async iteration and bridging
- Checkpoint — Concurrency
← concurrent.futures — executors, futures, map and as_completed · asyncio basics — coroutines, await, tasks and gather →