Python JSON: loads vs load, dumps Options and Nested Data
Python's json module maps JSON onto dicts, lists, strings, numbers, booleans and None. loads vs load, pretty printing with sorted keys, and walking nested data.
- Course: Python study plan
- Module: Dictionaries and sets
- Kind: Lesson
- Reading time: 14 min
- Runtime: CPython 3.11
How do I parse JSON in Python?
Parse JSON in Python with the standard json module: json.loads(text) turns a JSON string into Python objects and json.load(f) reads from a file object. Objects become dicts with string keys, arrays become lists, true and false become True and False, and null becomes None. json.dumps and json.dump go the other way.
Lesson
Real data is nested: a list of records, each a dict, some values themselves lists of dicts. JSON is the text form of exactly that shape — objects, arrays, strings, numbers, booleans and null — and Python's json module maps it onto dicts, lists, str, int/float, bool and None with no loss. This lesson covers walking nested structures safely, json.loads/dumps and their options, the round trip and what it does not preserve, deterministic output with sort_keys, pprint, and the recursive walk that handles a document of unknown depth.
The mapping
| JSON | Python |
|---|---|
object {"k": v} | dict (keys are always strings) |
array […] | list |
| string | str |
| number | int if it has no fraction or exponent, else float |
true / false | True / False |
null | None |
import json
doc = json.loads('{"name": "ada", "langs": ["python", "c"], "age": 36, "active": true, "boss": null}')
doc["langs"][0] # 'python'
doc["boss"] is None # True
json.dumps(doc) # '{"name": "ada", "langs": ["python", "c"], "age": 36, "active": true, "boss": null}'
loads/dumps work on strings; load/dump on file objects (json.load(sys.stdin) reads a whole document from standard input). Tuples serialise as arrays and come back as lists; sets, dates, Decimal and class instances have no JSON form and raise TypeError unless you supply default= (Module 14).
dumps options
json.dumps(doc, indent=2) # pretty-printed, nested by two spaces
json.dumps(doc, sort_keys=True) # keys in sorted order — deterministic output
json.dumps(doc, separators=(",", ":")) # compact: no spaces
json.dumps("café", ensure_ascii=False) # '"café"' rather than '"caf\\u00e9"'
sort_keys=True is the option to remember for judged output and for comparing documents: two equal dicts built in different orders serialise identically. indent with sort_keys is the standard "show me this structure" combination; pprint.pprint(doc) does the same for any Python object (sorting dict keys by default) without JSON's restrictions.
Walking with defaults
doc.get("address", {}).get("city") # None instead of KeyError at either level
langs = doc.get("langs") or [] # missing or null -> empty list
first = doc["langs"][0] if doc.get("langs") else None
Chained get with an empty default of the right type is the safe navigation idiom. When a missing field is a bug rather than an option, index directly and let KeyError say so. match (Module 3) is the cleanest way to validate a document's shape when several shapes are possible:
match event:
case {"type": "click", "pos": [int(x), int(y)]}: ...
case {"type": "key", "key": str(k)}: ...
case _: raise ValueError("unknown event")
Recursive walks
A document of unknown depth is walked with a recursive function that dispatches on the type of the current node:
def walk(node, path=""):
if isinstance(node, dict):
for key, value in node.items():
yield from walk(value, f"{path}.{key}" if path else key)
elif isinstance(node, list):
for i, value in enumerate(node):
yield from walk(value, f"{path}[{i}]")
else:
yield path, node
for path, leaf in walk(doc):
print(path, "=", leaf)
name = ada
langs[0] = python
langs[1] = c
age = 36
active = True
boss = None
The same skeleton counts leaves, finds every value under a given key, computes depth, or rewrites values in place (return a new structure rather than mutating while walking). yield from (Module 11) makes the recursion produce a flat stream.
Building nested structures
report = {"total": 0, "items": []}
for name, qty in rows:
report["items"].append({"name": name, "qty": qty})
report["total"] += qty
by_dept = defaultdict(list) # then
doc = {"departments": [{"name": d, "staff": sorted(v)} for d, v in sorted(by_dept.items())]}
Build dicts and lists as the data dictates, then serialise once. Sorting keys and lists before dumps is what makes the output stable when the source order was not.
The round trip and its limits
json.loads(json.dumps(x)) == x holds for dicts with string keys, lists, strings, ints, floats (within float precision), booleans and None. It does not hold for: tuples (become lists), non-string keys ({1: "a"} becomes {"1": "a"}), float("nan")/inf (emitted as NaN/Infinity, which is not valid JSON, unless allow_nan=False raises), and any custom object. When keys are integers, convert on the way back ({int(k): v for k, v in loaded.items()}).
Parse errors raise json.JSONDecodeError (a ValueError) with the line and column: trailing commas, single quotes and comments are all invalid JSON, however common they are in JavaScript.
Pitfalls
json.loadson a file object orjson.loadon a string (swap them).- Expecting tuples, sets or int keys to survive the round trip.
- Printing a nested structure without
sort_keysand comparing against expected output built in another order. doc["a"]["b"]on a document that may lack"a"; chaingets or catchKeyError.- Mutating a list while walking it recursively.
- Single quotes or trailing commas in hand-written JSON.
Key takeaways
- JSON objects, arrays, strings, numbers, booleans and null map to dict, list, str, int/float, bool and None; keys are always strings.
loads/dumpsfor strings,load/dumpfor files;indent,sort_keys=Trueandensure_ascii=Falsecontrol the text.- Navigate optional fields with chained
getand typed defaults; validate shapes withmatch. - Walk unknown depth with a recursive function dispatching on
dict/list/ leaf. - The round trip loses tuples and non-string keys; sort keys and lists for deterministic output.
Common questions
What is the difference between json.load and json.loads?
json.loads parses a string — the s stands for string — while json.load reads from a file object such as an open file or sys.stdin. Likewise, json.dumps returns a string and json.dump writes to a file.
How do I pretty-print JSON in Python?
Pass indent=2 to json.dumps for output nested by two spaces, and add sort_keys=True so that equal dicts built in different orders print identically. For any Python object, pprint.pprint does the same job without JSON's restrictions.
How do I safely access a nested dictionary key in Python?
Chain get calls with an empty default of the right type: doc.get("address", {}).get("city") returns None instead of raising KeyError at either level. When a missing field would be a bug, index directly and let the KeyError report it.
Why does json.dumps turn tuples into lists and int keys into strings?
JSON has only arrays and objects with string keys, so a tuple is written as an array and read back as a list, and {1: "a"} comes back as {"1": "a"}. Convert on the way back, for example with {int(k): v for k, v in loaded.items()}.
What does TypeError: Object of type set is not JSON serializable mean?
Sets, dates, Decimal values and class instances have no JSON form, so json.dumps raises TypeError for them. Convert the value first, such as sorted(s) for a set, or pass a default= function that turns unknown objects into something JSON can hold.
Exercises
Walk the document
Read one JSON document from standard input with json.load(sys.stdin) and print every leaf value with its path, using a recursive walk: object keys join with ., array indexes appear as [i]. Print booleans and null the Python way (True, None). Finish with the number of leaves.
Input: a JSON document. Output: <path> = <value> per leaf in document order, then leaves: <n>.
{"name": "ada", "langs": ["python", "c"], "active": true, "boss": null}
prints
name = ada
langs[0] = python
langs[1] = c
active = True
boss = None
leaves: 5Canonical JSON
Read a JSON document and print it in canonical form — json.dumps with sorted keys and compact separators — then the number of top-level keys (0 if the document is not an object) and the maximum nesting depth, where a scalar has depth 0 and each enclosing object or array adds 1.
Input: a JSON document. Output: the canonical text, then keys: <n>, then depth: <d>.
{"b": [1, {"z": null, "a": 2}], "a": "x"}
prints
{"a":"x","b":[1,{"a":2,"z":null}]}
keys: 2
depth: 3In this module: Dictionaries and sets
- Dictionaries — the mapping at the centre of Python
- Counting and grouping — Counter, defaultdict and the accumulation idioms
- Sets — membership, deduplication and set algebra
- Hashing and keys — what makes an object usable in a dict or set
- Nested data and JSON (this lesson)
- Choosing a collection — the complexity table and heapq
- Checkpoint — Dictionaries and sets
← Hashing and keys — what makes an object usable in a dict or set · Choosing a collection — the complexity table and heapq →