Python ThreadPoolExecutor and concurrent.futures Explained
The concurrent.futures module runs calls on thread or process pools. Submit vs map, ordered results vs completion order, exceptions, timeouts and pool size.
- Course: Python study plan
- Module: Concurrency — threads, processes and asyncio
- Kind: Lesson
- Reading time: 13 min
- Runtime: CPython 3.11
What is ThreadPoolExecutor in Python?
ThreadPoolExecutor is the thread pool in Python's concurrent.futures module: submit(fn, *args) schedules a call on a free worker and returns a Future, and map(fn, items) returns the results in input order. Used as a with block, it waits for every job before exiting. ProcessPoolExecutor has the same interface but runs the calls in separate processes, for CPU-bound work.
Lesson
concurrent.futures is the interface most concurrent Python code should use: an executor runs callables on a pool of threads or processes, each call returns a future that will hold the result, and the same three methods — submit, map, shutdown — work for both pools. The result of a job is fetched with future.result(), which re-raises any exception the job raised, and map hands results back in submission order — the property that makes threaded output deterministic. This lesson covers the two executors, submit versus map, as_completed and wait, exceptions and timeouts, pool sizing, and when to switch from threads to processes.
The executor and the future
from concurrent.futures import ThreadPoolExecutor
def fetch(url):
return len(download(url))
with ThreadPoolExecutor(max_workers=4) as pool: # shutdown(wait=True) on exit
future = pool.submit(fetch, "https://a") # starts as soon as a worker is free
...
size = future.result() # blocks until done; re-raises if fetch raised
submit(fn, *args, **kwargs) schedules one call and returns a Future immediately. result(timeout=None) waits for and returns the value; exception() returns the exception instead of raising; done(), running(), cancel() inspect and control it; add_done_callback(fn) runs fn(future) on completion. The with block waits for every submitted job before exiting, so results are always complete after it.
map: ordered results
with ThreadPoolExecutor(max_workers=4) as pool:
sizes = list(pool.map(fetch, urls)) # results in the order of urls, whatever finished first
map(fn, *iterables, timeout=None, chunksize=1) submits every element and yields results in input order, blocking on each in turn. It is the one-line replacement for "start N threads, join them, read the slots" and the tool to reach for first. Its limitation: an exception from any element is raised when its position is reached, and iteration stops there — for per-job error handling, use submit and inspect each future.
as_completed and wait
from concurrent.futures import as_completed, wait, FIRST_COMPLETED
futures = {pool.submit(fetch, url): url for url in urls}
for future in as_completed(futures): # yields futures as they FINISH — arbitrary order
url = futures[future]
try:
print(url, future.result())
except Exception as e:
print(url, "failed:", e)
done, pending = wait(futures, timeout=5, return_when=FIRST_COMPLETED)
as_completed is for showing progress or acting on results as they arrive; because its order is the completion order, anything it produces must be sorted before it is printed as an answer. wait blocks until all, or the first, or the first exception, with a timeout, and returns the two sets. Mapping futures back to their inputs with a dict, as above, is the standard idiom.
Processes
from concurrent.futures import ProcessPoolExecutor
def cpu_heavy(n):
return sum(i * i for i in range(n))
if __name__ == "__main__": # required: workers import the main module
with ProcessPoolExecutor() as pool: # max_workers defaults to the CPU count
totals = list(pool.map(cpu_heavy, [10**6] * 4, chunksize=1))
Same interface, separate interpreters, real parallelism for CPU-bound functions. The constraints: the function must be importable by name (a module-level def, not a lambda or a nested function), arguments and results must be picklable, the main module must guard its entry point (workers re-import it), and each task carries a fixed cost of pickling and process communication — batch small tasks with chunksize or they will be slower than a plain loop.
Exceptions and timeouts
A job that raises does not crash the pool; the exception is stored in its future and re-raised by result() (or returned by exception()). result(timeout=2) raises TimeoutError if the job is not done in time — the job itself keeps running; futures cannot interrupt a running callable, only cancel() one that has not started. Design long jobs to check a stop flag or an Event themselves.
Sizing the pool
For threads doing I/O, more workers than cores is normal — 10 to 50 for network calls, bounded by what the remote side tolerates; the default is min(32, cpu_count + 4). For processes doing CPU work, the CPU count (the default) is the ceiling that helps. Too many workers on a rate-limited API produces failures, not speed; a Semaphore inside the job or a smaller pool is the fix. Measure with perf_counter around the whole with block.
A pattern for deterministic output
def process(item):
return item, expensive(item) # return the key with the result
with ThreadPoolExecutor() as pool:
results = dict(pool.map(process, items)) # or sorted(...) — order fixed by the data, not the schedule
for key in sorted(results):
print(key, results[key])
Return the identity of the work with its result, collect everything inside the with, then print in an order defined by the data. This is the shape every judged exercise in this module takes.
Pitfalls
- Printing inside the job.
- Reading
as_completedresults as if they were ordered. - A lambda or nested function submitted to a process pool.
- No
if __name__ == "__main__":around a process pool on Windows and macOS (recursive process spawning). mapswallowing per-item errors into one raise at the failing position.- A pool of 100 threads against an API that allows 5 concurrent requests.
Key takeaways
ThreadPoolExecutor/ProcessPoolExecutorwith the same interface:submit→Future,map→ ordered results,withwaits for everything.future.result()blocks and re-raises;exception(),done(),cancel(),add_done_callbackinspect and control.mappreserves input order — the deterministic tool;as_completedyields by completion and needs sorting;waithandles first/all/timeout.- Processes need importable, picklable work and a main guard; batch with
chunksize. - Return the key with the result, collect inside the
with, print in data order.
Common questions
What is the difference between executor.map and as_completed?
map yields results in the order the inputs were submitted, whatever finished first, which makes output deterministic. as_completed yields futures in the order they finish, which suits progress reporting; sort its results before printing them as an answer.
How do exceptions work in a ThreadPoolExecutor?
A job that raises does not crash the pool. The exception is stored in its Future and re-raised when you call future.result(), or returned by future.exception(). With map, the exception is raised when iteration reaches that item, and iteration stops there.
Can a Future timeout stop a running job?
No. future.result(timeout=2) raises TimeoutError if the job has not finished, but the job keeps running; cancel() only works on a job that has not started. Long jobs must check a stop flag or an Event themselves.
How many workers should a Python thread pool have?
For I/O-bound threads, more workers than cores is normal: the default is min(32, cpu_count + 4), and 10 to 50 suits network calls, bounded by what the remote side tolerates. For CPU-bound process pools, the CPU count is the ceiling that helps.
Why does ProcessPoolExecutor need a main guard?
Under the spawn start method, the default on Windows and macOS, each worker imports the main module to find the function, so without if __name__ == '__main__': every worker would run the pool-creating code again. The function must also be defined at module level, and its arguments must be picklable.
Exercises
Ordered map over words
Read one line of words. With a ThreadPoolExecutor(max_workers=3), use map to compute for each word a tuple (word, len(word), vowels) where vowels counts the letters in aeiou. Print <word> len=<n> vowels=<v> in input order — map preserves it — then most vowels <word> (ties go to the earlier word).
Input: one line of words. Output: one line per word, then the winner.
queue thread future
prints
queue len=5 vowels=4
thread len=6 vowels=2
future len=6 vowels=3
most vowels queuePrimes in a process pool
Write a module-level prime_sum(n) that returns the sum of the primes below n using a sieve. Read one line of integers; under if __name__ == "__main__": run them through a ProcessPoolExecutor(max_workers=2) with map and print primes below <n> sum to <s> in input order.
Input: one line of integers. Output: one line per integer.
10 20 100
prints
primes below 10 sum to 17
primes below 20 sum to 77
primes below 100 sum to 1060In this module: Concurrency — threads, processes and asyncio
- The GIL and the three models — threads, processes, asyncio
- Threads — Thread, Lock, Event, Queue and the race you must see once
- concurrent.futures — executors, futures, map and as_completed (this lesson)
- multiprocessing — Process, Pool, pickling, queues and shared state
- asyncio basics — coroutines, await, tasks and gather
- asyncio patterns — TaskGroup, queues, semaphores, async iteration and bridging
- Checkpoint — Concurrency
← Threads — Thread, Lock, Event, Queue and the race you must see once · multiprocessing — Process, Pool, pickling, queues and shared state →