Python Generator Expressions vs List Comprehensions
A Python generator expression is a lazy, one-shot comprehension in parentheses. When it beats a list comprehension and how any, all and next stop early.
- Course: Python study plan
- Module: Iterators, generators and itertools
- Kind: Lesson
- Reading time: 12 min
- Runtime: CPython 3.11
What is a generator expression in Python?
A generator expression is a comprehension written in parentheses, such as (x * x for x in xs), that produces a generator instead of a list. Values are computed only as a consumer asks for them, so memory stays constant however long the input is. It is one-shot, and it cannot be indexed, sliced or passed to len.
Lesson
A generator expression is a comprehension in parentheses that produces a generator instead of a list: (x * x for x in xs). It is the same lazy machinery as a generator function, written inline for the one-line cases — the argument to sum, max, any, all, "".join, sorted, set, dict, or a for loop. Choosing it over a list comprehension is a memory decision (O(1) instead of O(n)) and sometimes a time decision (short-circuit consumers stop early). This lesson covers the syntax and its one-shot nature, when it beats a list comprehension and when it does not, the short-circuit consumers, the next(gen, default) idiom, and the scoping detail that makes the first for clause evaluate eagerly.
The syntax
squares = (x * x for x in range(5)) # a generator object; nothing computed yet
type(squares) # <class 'generator'>
list(squares) # [0, 1, 4, 9, 16]
list(squares) # [] — one-shot, like every generator
total = sum(x * x for x in range(5)) # the parentheses of the call are enough
big = sum((x for x in xs), start=0) # with a second argument, the generator needs its own parentheses
Everything a list comprehension can express — a filtering if, nested for clauses, a conditional expression — works identically. The only difference is that the result is produced lazily and cannot be indexed, sliced, measured with len or iterated twice.
Generator versus list comprehension
| Situation | Use |
|---|---|
Feeding a single consumer once (sum, max, any, join, for) | generator expression |
The values are needed more than once, indexed, sliced or len-ed | list comprehension |
| The data is large or the source is itself lazy (a file, another generator) | generator expression |
The consumer may stop early (any, all, next, in) | generator expression — work stops at the answer |
| You want to print or debug the intermediate values | list comprehension (a generator shows as <generator object …>) |
sum([x * x for x in range(10 ** 6)]) # builds a million-element list, then sums it
sum(x * x for x in range(10 ** 6)) # sums as it goes; constant memory
For sum and friends the generator is usually also faster, because no list is allocated and grown. For sorted and set, which must see everything, a list comprehension and a generator expression cost about the same — the generator saves memory, not time.
Short-circuit consumers
any, all, in (on a generator) and next stop at the first value that decides the answer, and a generator expression feeds them exactly as many values as they need:
any(x < 0 for x in xs) # stops at the first negative
all(len(w) < 10 for w in words) # stops at the first long word
next((x for x in xs if x % 7 == 0), None) # the first multiple of 7, or None
first_long = next((w for w in words if len(w) > 8), "")
With a list comprehension, any([x < 0 for x in xs]) computes every element before any looks at the first. The generator version is the idiom for "does one exist?" and "give me the first that…", and next(gen, default) is how "first that…" handles the case where nothing matches.
In other calls
", ".join(str(n) for n in nums) # join needs strings; the generator converts lazily
max(words, key=len) # not a generator: a key function
dict((k, len(k)) for k in keys) # a dict from pairs; the dict comprehension {k: len(k) for k in keys} is clearer
set(w.lower() for w in words) # a set; {w.lower() for w in words} is clearer
sorted(x for x in xs if x) # fine: sorted materialises anyway
tuple(x for x in xs) # there is no tuple comprehension; this is it
When a set, dict or list is the goal, use the matching comprehension; when the goal is a computation over the values, pass a generator expression to it.
The eager first clause
The iterable of the first for clause is evaluated immediately, when the generator expression is created; everything else is evaluated lazily as values are pulled:
def source():
print("source evaluated")
return [1, 2, 3]
g = (x for x in source()) # prints "source evaluated" now
values = list(g) # the loop body runs here
This means an error in the first iterable surfaces at creation, and that later clauses (for y in f(x)) run lazily. It also means a generator expression captures its outer variables by reference, like a lambda — the late-binding trap of Module 4 applies to a generator created in a loop and consumed later.
Nesting and readability
A generator expression inside another call inside another call is hard to read; give the inner one a name:
lengths = (len(line) for line in lines)
print(f"longest: {max(lengths)}")
Two clauses is fine; three, or a conditional expression plus a filter, wants a generator function with a name and a docstring.
Pitfalls
- Iterating a generator expression twice.
len(gen),gen[0],gen[:3]—TypeError.sum([... for ...])when a generator would do.any([...])losing the short-circuit.- Printing a generator expression and seeing
<generator object>. - A generator expression capturing a loop variable that changes before it is consumed.
Key takeaways
(expr for x in it if cond)is a lazy, one-shot comprehension; the call's parentheses suffice when it is the only argument.- Use it to feed
sum,max,any,all,join,nextandfor; use a list comprehension when values are reused, indexed or measured. any/all/nextshort-circuit, so a generator does only the work the answer needs;next(gen, default)is "first that…".- The first
foriterable is evaluated at creation; the rest lazily; outer names are captured by reference. - Prefer the set/dict comprehension when a set or dict is the goal; name a generator expression when it gets long.
Common questions
What is the difference between a generator expression and a list comprehension?
A list comprehension in square brackets builds the whole list at once, holding every element in memory; a generator expression in parentheses produces values lazily, one at a time. Use the generator to feed a single consumer such as sum, any or join, and the list when values are reused, indexed or measured.
Is sum() faster with a generator or a list comprehension?
The generator usually wins and always uses less memory: sum(x * x for x in xs) adds values as they are produced, while sum([x * x for x in xs]) first allocates and grows a full list. For sorted and set, which must see every value anyway, the two cost about the same time.
How do I get the first item that matches a condition in Python?
Use next() with a generator expression and a default: next((x for x in xs if x % 7 == 0), None). The generator stops at the first match, so no further work is done, and the default is returned instead of StopIteration when nothing matches.
Is there a tuple comprehension in Python?
No. (x for x in xs) in parentheses is a generator expression, not a tuple. To build a tuple, pass a generator expression to the constructor: tuple(x for x in xs). Sets and dicts have their own comprehensions, written with braces.
Exercises
Short-circuit drills
Read a threshold and a line of integers. With generator expressions only (no lists), print: whether any value exceeds the threshold; whether all values are positive; the first value exceeding the threshold or none (using next with a default); how many exceed it (sum); and the values that exceed it joined by commas (join over a generator of strings).
Input: the threshold, then a line of integers. Output: any <bool>, all-positive <bool>, first <v|none>, count <n>, above <a,b,…> (or above none).
5
3 8 1 9
prints
any True
all-positive True
first 8
count 2
above 8,9Lazy versus eager, traced
compute(i) prints compute <i> and returns i * i. Read n and target. First evaluate any(compute(i) > target for i in range(n)) — a generator expression, which stops at the first hit — then any([compute(i) > target for i in range(n)]) — a list comprehension, which computes everything first. Print the trace lines as they happen and the two results as lazy <bool> and eager <bool>.
Input: n target. Output: the trace and the two result lines.
4 3
prints
compute 0
compute 1
compute 2
lazy True
compute 0
compute 1
compute 2
compute 3
eager TrueIn this module: Iterators, generators and itertools
- The iteration protocol — iter, next and StopIteration
- Generators — functions that yield
- Generator expressions — lazy comprehensions (this lesson)
- itertools — the iterator toolkit
- functools and operator — the function toolkit
- Lazy pipelines — a worked log-processing example
- Checkpoint — Iterators, generators and itertools
← Generators — functions that yield · itertools — the iterator toolkit →