Python Unit Testing Best Practices: Arrange, Act, Assert
What to test in Python and how: unit vs integration vs end-to-end tests, arrange-act-assert, edge cases, flaky tests, and designing code for testability.
- Course: Python study plan
- Module: Testing and debugging
- Kind: Lesson
- Reading time: 13 min
- Runtime: CPython 3.11
What makes a good unit test in Python?
A good unit test in Python checks one behaviour of one function or class, with no I/O, in three parts: arrange the inputs, act with one call, assert on the outcome. It is named for the behaviour it checks, runs in milliseconds, sets up its own state and gives the same result every run, so a failure explains itself and nobody learns to ignore it.
Lesson
A test is a program that runs your code with known inputs and checks the outputs, so that a change which breaks behaviour is caught by a machine instead of a user. That definition already settles most arguments: tests exist to catch regressions and to pin down what the code promises, not to prove correctness or to reach a coverage number. This lesson covers the kinds of test and where each earns its place, the arrange–act–assert shape, what is worth testing, the properties that make a suite trustworthy, and the design moves — pure functions, injected dependencies, fixed clocks and seeds — that make code testable in the first place.
Kinds of test
| Kind | Scope | Speed | Count |
|---|---|---|---|
| Unit | one function or class, no I/O | microseconds to milliseconds | hundreds or thousands |
| Integration | two or more real parts together: code and a database, a parser and a file | milliseconds to seconds | tens |
| End-to-end | the whole system through its real interface: an HTTP request, a CLI invocation | seconds | a handful |
The pyramid — many unit tests, fewer integration tests, few end-to-end — follows from cost: a unit test runs in isolation and tells you which line is wrong; an end-to-end test proves the system works but says little about where it failed and takes a thousand times longer. A regression test is any test written to reproduce a bug before fixing it; it belongs at the lowest level that reproduces the problem.
Arrange, act, assert
def test_discount_applies_over_threshold():
cart = Cart([Item("book", 40), Item("pen", 70)]) # arrange: the inputs
total = cart.total(discount_over=100, rate=0.1) # act: one call
assert total == 99.0 # assert: one behaviour
Every test has these three parts, in this order, ideally with a blank line between them. One behaviour per test, named for the behaviour (test_discount_applies_over_threshold, not test_total_2), so that a failure's name is already a diagnosis. Several assertions are fine when they describe one behaviour; a test that checks five unrelated things stops at the first failure and hides the other four.
What to test
- Behaviour, not implementation. Assert on what the function returns or does to the world, not on which helpers it called; a test coupled to the implementation fails on every refactor and passes on every bug that keeps the same shape.
- The edges. Empty input, one element, the boundary value (exactly 100 when the rule is "over 100"), negative, zero, the largest sensible size, Unicode, whitespace.
- The errors. The invalid inputs raise the documented exception with a useful message; the recoverable failures are recovered.
- The bug you just fixed. A regression test that failed before the fix and passes after.
- Not private helpers, trivial getters, or the standard library.
Properties of a good suite
Fast — a suite that takes minutes is run rarely. Isolated — each test sets up its own state and leaves nothing behind, so order does not matter and one failure does not cascade. Deterministic — the same result every run: no dependence on the clock, randomness, network, the file system's contents, or hash order. Readable — a failing test explains itself from its name and its assertion message. Independent of each other — no test relies on another having run. A flaky test (passes sometimes) is worse than no test: it trains people to ignore red.
Designing for testability
The code that is easy to test is code whose inputs and outputs are explicit:
# hard to test: reads the clock, the environment and the network inside
def report():
today = date.today()
rows = requests.get(URL).json()
...
# easy to test: the world is passed in
def report(rows, today):
...
def main():
print(report(fetch_rows(), date.today())) # the impure shell, thin
The pattern is functional core, imperative shell: pure functions that compute from arguments, wrapped by a thin layer that does the I/O. Dependencies a function needs — the clock, a random generator, a sender, a store — are parameters with sensible defaults (now=None then now = now or datetime.now()), so a test passes a fixed value and the production caller passes nothing. random.Random(seed) instead of the module functions makes randomness reproducible. A class that takes its collaborators in __init__ can be given fakes.
Doubles, briefly
When a dependency is slow, unavailable or nondeterministic, tests replace it with a double: a stub returns canned answers, a fake is a working lightweight implementation (an in-memory dict for a database), a spy records what was called, a mock does both and can assert on calls. Lesson 4 builds them with unittest.mock; the design above is what makes them pluggable.
Coverage and TDD
Coverage (coverage run -m pytest, then coverage report) shows which lines tests executed. Untested lines are a signal; 100 % is not a goal — a line can execute without its behaviour being checked. Test-driven development writes the failing test first, then the least code that passes, then refactors; it is a design technique as much as a testing one, because a test written first forces the interface to be usable. Neither is mandatory; both are tools.
Pitfalls
- Tests that assert on implementation details (which method was called, in what order) rather than outcomes.
- A test with no assertion (it "runs without crashing").
- Hidden dependence on the clock, randomness or test order.
- One giant test per module.
- Mocking so much that the test checks the mocks, not the code.
- Skipping the regression test because "the fix is obvious".
Key takeaways
- Tests catch regressions and pin promises; unit tests are many and fast, integration tests fewer, end-to-end tests few.
- Arrange, act, assert; one behaviour per test, named for it.
- Test behaviour, edges and errors, and every fixed bug; not private helpers.
- A trustworthy suite is fast, isolated, deterministic and readable — flaky tests are worse than none.
- Make code testable by passing the world in: pure core, thin impure shell, injected clock, seed and collaborators.
Common questions
What is the arrange-act-assert pattern?
The three parts of every test, in order: arrange the inputs, act with one call to the code under test, assert on the outcome. Checking one behaviour per test means a failure points at one thing, whereas a test of five unrelated things stops at the first failure and hides the rest.
What is the difference between unit, integration and end-to-end tests?
A unit test checks one function or class with no I/O, in milliseconds; an integration test runs real parts together, such as code and a database; an end-to-end test drives the whole system through its real interface. The test pyramid is many unit tests, fewer integration tests and a handful end to end.
What should I not unit test?
Private helpers, trivial getters and the standard library. Test behaviour through the public interface, meaning what a function returns or changes, plus the edge cases, the documented errors and every bug you fix, rather than which helpers the code happened to call.
How do I make Python code easier to test?
Pass the world in. Keep a functional core of pure functions that compute from their arguments inside a thin shell that does the I/O, and make the clock, the random generator and collaborators parameters with defaults, so a test can supply a fixed time, a seeded random.Random or a fake.
Why is a flaky test worse than no test?
A test that passes only sometimes trains people to ignore red, so real regressions get ignored too. Flakiness usually comes from hidden dependence on the clock, randomness, the network, hash order or the order in which tests run.
Is 100% code coverage worth aiming for?
No. Coverage shows which lines ran, not whether their behaviour was checked, so untested lines are a useful signal but a perfect score is not a goal. coverage run -m pytest followed by coverage report measures it.
Exercises
A table of cases
Implement parse_duration(text) returning seconds for strings made of optional <n>h, <n>m and <n>s parts in that order (1h30m, 45s, 2h), raising ValueError for anything else. Then read lines <input> <expected> until EOF, where expected is an integer or the word error, and run each through the function: print ok <input> -> <got> when the result (or error when it raised) matches, else FAIL <input>: expected <expected>, got <got>. Finish with passed <x> of <y>.
Input: one case per line. Output: one line per case, then the summary.
1h 3600
30m 1800
1h30m 5400
45s 45
2h 7000
x error
prints
ok 1h -> 3600
ok 30m -> 1800
ok 1h30m -> 5400
ok 45s -> 45
FAIL 2h: expected 7000, got 7200
ok x -> error
passed 5 of 6Inject the clock
Write status(due, today) taking two datetime.date values and returning overdue by <n> day(s) when due is before today, due today when equal, and due in <n> day(s) otherwise — with day for 1 and days otherwise. Read today as an ISO date on the first line, then lines <name> <YYYY-MM-DD> until EOF; print <name>: <status> per line, then overdue <count>. The date is a parameter, never date.today(), so the program is testable.
Input: the date, then one item per line. Output: one line per item, then the count.
2024-03-10
tax 2024-03-01
rent 2024-03-10
gym 2024-03-15
prints
tax: overdue by 9 days
rent: due today
gym: due in 5 days
overdue 1In this module: Testing and debugging
- The testing mindset — what to test, how to arrange it, and designing for testability (this lesson)
- unittest — TestCase, assertions, fixtures, subTest and running suites
- pytest in outline — plain asserts, fixtures, parametrize and the command line
- Mocking and test doubles — Mock, patch, side_effect and where to patch
- Debugging — reading tracebacks, logging, pdb and the method
- Checkpoint — Testing and debugging
← Checkpoint — Concurrency · unittest — TestCase, assertions, fixtures, subTest and running suites →