Python Counter and defaultdict: Counting and Grouping
Count with collections.Counter and group with defaultdict. How Counter orders ties, nested defaultdicts, inverting a dict, and why groupby needs sorted input.
- Course: Python study plan
- Module: Dictionaries and sets
- Kind: Lesson
- Reading time: 13 min
- Runtime: CPython 3.11
How do I count occurrences of items in Python?
Use collections.Counter: Counter(words) builds a dict subclass mapping each item to how often it appears, and a missing key reads as 0 instead of raising KeyError. most_common(n) returns the n commonest pairs, highest count first, with ties in first-seen order. On a plain dict the same count is counts[w] = counts.get(w, 0) + 1.
Lesson
Half of all dictionary code does one of two things: counts how often each key occurs, or collects the items that share a key into a list. Both have a plain-dict spelling, a setdefault spelling, and a collections type that says exactly what is meant — Counter for counting, defaultdict for grouping. This lesson gives all three for each job, the Counter API (most_common, arithmetic, elements), defaultdict with different factories including nested ones, inverting a mapping, and the reason itertools.groupby is not the grouping tool it sounds like.
Counting
counts = {}
for word in words:
if word in counts: # plain dict, verbose
counts[word] += 1
else:
counts[word] = 1
counts[word] = counts.get(word, 0) + 1 # plain dict, one line
from collections import Counter
counts = Counter(words) # the type that means "count"
Counter is a dict subclass whose missing keys read as 0 (never KeyError) and which is built from any iterable or mapping:
c = Counter("mississippi")
c["s"] # 4
c["z"] # 0 — no KeyError
c.most_common(2) # [('i', 4), ('s', 4)] — ties keep first-seen order
c.total() # 11 (3.10)
c.update("miss") # add counts
c.subtract("ss") # subtract counts (can go negative)
list(c.elements()) # each key repeated by its count
Counter(a) + Counter(b) # add counts; & is min, | is max; - drops non-positive
sorted(c.items()) # deterministic order for output
most_common() with no argument returns everything sorted by count descending, with ties in insertion order — which is the deterministic ordering to print when the tie-break does not matter; when it does, sort explicitly: sorted(c.items(), key=lambda kv: (-kv[1], kv[0])) for count descending then key ascending.
Grouping
groups = {}
for name, dept in staff:
groups.setdefault(dept, []).append(name) # plain dict
from collections import defaultdict
groups = defaultdict(list)
for name, dept in staff:
groups[dept].append(name) # the type that means "group"
defaultdict(factory) calls factory() to create the value for a missing key on first access and stores it; afterwards it is an ordinary dict. The factory is any zero-argument callable: list, set, int (counting), dict, lambda: [0, 0], or a class. Nested groupings compose: defaultdict(lambda: defaultdict(int)) counts pairs as table[a][b] += 1.
Two behaviours to know. Reading a missing key creates it (groups["x"] inserts an empty list), so use in to test presence without side effects. And defaultdict prints as defaultdict(<class 'list'>, {...}); convert with dict(groups) for a clean display, and sort the keys for output.
Inverting and indexing
inverse = {v: k for k, v in d.items()} # one-to-one only: later keys overwrite
by_value = defaultdict(list) # many-to-one: group keys by value
for k, v in d.items():
by_value[v].append(k)
index = defaultdict(list) # a word index: word -> line numbers
for n, line in enumerate(lines, start=1):
for word in set(line.split()):
index[word].append(n)
The comprehension inverts a mapping whose values are unique; when several keys share a value, the grouping form keeps them all.
Accumulating other things
The same shape works for any per-key accumulation:
totals = defaultdict(float)
for name, amount in transactions:
totals[name] += amount
best = {}
for name, score in results:
best[name] = max(best.get(name, float("-inf")), score)
first_seen = {}
for i, key in enumerate(keys):
first_seen.setdefault(key, i) # keeps the first, ignores later
Counting with defaultdict(int) and with Counter are equivalent; Counter adds the API and reads as intent.
Counting pairs and combinations
A key can be a tuple, so co-occurrence counts need no nesting: Counter((a, b) for a, b in pairs) tallies each ordered pair, Counter(frozenset((a, b)) for a, b in pairs) each unordered one, and Counter(zip(words, words[1:])) counts adjacent word pairs — bigrams — in one line. most_common then answers "which pair is commonest" directly.
itertools.groupby is not this
itertools.groupby(iterable, key) groups consecutive elements with equal keys — it is a run-length tool, and on unsorted input it produces one group per run, not per key. It groups whole data only after a sorted(..., key=key), and even then yields iterators that must be consumed immediately. For "group all items by key", defaultdict(list) is the tool; groupby is for runs, or for already-sorted data streamed in one pass (Module 11).
Printing counts and groups
Output must be deterministic, and a dict's order is insertion order — fine when the input order is meaningful, but usually you want sorted keys or descending counts:
for key in sorted(groups):
print(key, ", ".join(groups[key]))
for key, n in sorted(counts.items(), key=lambda kv: (-kv[1], kv[0])):
print(f"{key}: {n}")
Never print a set or iterate one into output (lesson 3); a Counter's most_common is fine because its order is defined.
Pitfalls
counts[k] += 1on a plain dict for a new key —KeyError; useget,defaultdict(int)orCounter.- Testing
if key in groupsafter agroups[key]read — the read already created it. {v: k …}inverting a many-to-one mapping and losing keys.groupbyon unsorted data.- Printing a
defaultdictor aCounterdirectly and getting the class name in the output. - Relying on
most_commontie order when the spec says ties break alphabetically.
Key takeaways
- Count with
Counter(iterable); missing keys read as 0;most_common,update,+/-/&/|,total,elements. - Group with
defaultdict(list); any zero-argument factory works, including nested defaultdicts. get(k, 0) + 1andsetdefault(k, []).append(x)are the plain-dict spellings.- Invert with a comprehension when values are unique, with grouping when they are not.
itertools.groupbygroups consecutive runs only; sort first or usedefaultdict.
Common questions
What is defaultdict in Python?
collections.defaultdict(factory) is a dict that calls factory() to create the value the first time a missing key is accessed, and stores it. defaultdict(list) groups items with groups[key].append(x) and defaultdict(int) counts; any zero-argument callable works as the factory.
What is the difference between Counter and defaultdict(int)?
Both count without a KeyError, but Counter adds a counting API: most_common, update, subtract, total, elements, and arithmetic such as + and &. It also reads a missing key as 0 without inserting it, whereas reading a missing key from a defaultdict creates the entry.
How do I group a list of items by key in Python?
Create groups = defaultdict(list) and append each item under its key: groups[dept].append(name). The plain-dict spelling is groups.setdefault(dept, []).append(name). Sort the keys before printing, since a dict iterates in insertion order.
Why does itertools.groupby not group all my items?
itertools.groupby groups only consecutive elements with equal keys, so unsorted input gives one group per run rather than one per key. Sort by the same key first, or use defaultdict(list), which collects every item for a key wherever it appears.
How do I invert a dictionary in Python?
When the values are unique, {v: k for k, v in d.items()} swaps keys and values. When several keys share a value, later keys overwrite earlier ones, so group instead: append each key to by_value[v] in a defaultdict(list).
Exercises
Word frequencies
Read text until the end of input and count words (case-folded, with leading and trailing punctuation stripped; empty results ignored) with a Counter. Read k from the first line. Print the k most frequent words as word count, ordering ties alphabetically (so sort explicitly rather than relying on most_common), then the number of words that occur exactly once.
Input: k, then lines of text. Output: up to k lines, then once: <n>.
2
the cat and the hat. The end, and... fin
prints
the 3
and 2
once: 4Group by department
Read lines name dept until the end of input and group the names by department with a defaultdict(list). Print each department in alphabetical order with its names sorted and comma-separated, then the largest department (ties broken alphabetically).
Input: lines. Output: <dept>: <names> per department, then largest: <dept> (or largest: none).
ada eng
bob ops
cy eng
prints
eng: ada,cy
ops: bob
largest: engIn this module: Dictionaries and sets
- Dictionaries — the mapping at the centre of Python
- Counting and grouping — Counter, defaultdict and the accumulation idioms (this lesson)
- Sets — membership, deduplication and set algebra
- Hashing and keys — what makes an object usable in a dict or set
- Nested data and JSON
- Choosing a collection — the complexity table and heapq
- Checkpoint — Dictionaries and sets
← Dictionaries — the mapping at the centre of Python · Sets — membership, deduplication and set algebra →