Python JSON: loads vs load, dumps Options and Nested Data

Python's json module maps JSON onto dicts, lists, strings, numbers, booleans and None. loads vs load, pretty printing with sorted keys, and walking nested data.

  • Course: Python study plan
  • Module: Dictionaries and sets
  • Kind: Lesson
  • Reading time: 14 min
  • Runtime: CPython 3.11

How do I parse JSON in Python?

Parse JSON in Python with the standard json module: json.loads(text) turns a JSON string into Python objects and json.load(f) reads from a file object. Objects become dicts with string keys, arrays become lists, true and false become True and False, and null becomes None. json.dumps and json.dump go the other way.

Lesson

Real data is nested: a list of records, each a dict, some values themselves lists of dicts. JSON is the text form of exactly that shape — objects, arrays, strings, numbers, booleans and null — and Python's json module maps it onto dicts, lists, str, int/float, bool and None with no loss. This lesson covers walking nested structures safely, json.loads/dumps and their options, the round trip and what it does not preserve, deterministic output with sort_keys, pprint, and the recursive walk that handles a document of unknown depth.

The mapping

JSONPython
object {"k": v}dict (keys are always strings)
array […]list
stringstr
numberint if it has no fraction or exponent, else float
true / falseTrue / False
nullNone
import json
doc = json.loads('{"name": "ada", "langs": ["python", "c"], "age": 36, "active": true, "boss": null}')
doc["langs"][0]                    # 'python'
doc["boss"] is None                # True
json.dumps(doc)                    # '{"name": "ada", "langs": ["python", "c"], "age": 36, "active": true, "boss": null}'

loads/dumps work on strings; load/dump on file objects (json.load(sys.stdin) reads a whole document from standard input). Tuples serialise as arrays and come back as lists; sets, dates, Decimal and class instances have no JSON form and raise TypeError unless you supply default= (Module 14).

dumps options

json.dumps(doc, indent=2)                  # pretty-printed, nested by two spaces
json.dumps(doc, sort_keys=True)            # keys in sorted order — deterministic output
json.dumps(doc, separators=(",", ":"))     # compact: no spaces
json.dumps("café", ensure_ascii=False)     # '"café"' rather than '"caf\\u00e9"'

sort_keys=True is the option to remember for judged output and for comparing documents: two equal dicts built in different orders serialise identically. indent with sort_keys is the standard "show me this structure" combination; pprint.pprint(doc) does the same for any Python object (sorting dict keys by default) without JSON's restrictions.

Walking with defaults

doc.get("address", {}).get("city")                  # None instead of KeyError at either level
langs = doc.get("langs") or []                      # missing or null -> empty list
first = doc["langs"][0] if doc.get("langs") else None

Chained get with an empty default of the right type is the safe navigation idiom. When a missing field is a bug rather than an option, index directly and let KeyError say so. match (Module 3) is the cleanest way to validate a document's shape when several shapes are possible:

match event:
    case {"type": "click", "pos": [int(x), int(y)]}: ...
    case {"type": "key", "key": str(k)}: ...
    case _: raise ValueError("unknown event")

Recursive walks

A document of unknown depth is walked with a recursive function that dispatches on the type of the current node:

def walk(node, path=""):
    if isinstance(node, dict):
        for key, value in node.items():
            yield from walk(value, f"{path}.{key}" if path else key)
    elif isinstance(node, list):
        for i, value in enumerate(node):
            yield from walk(value, f"{path}[{i}]")
    else:
        yield path, node

for path, leaf in walk(doc):
    print(path, "=", leaf)
name = ada
langs[0] = python
langs[1] = c
age = 36
active = True
boss = None

The same skeleton counts leaves, finds every value under a given key, computes depth, or rewrites values in place (return a new structure rather than mutating while walking). yield from (Module 11) makes the recursion produce a flat stream.

Building nested structures

report = {"total": 0, "items": []}
for name, qty in rows:
    report["items"].append({"name": name, "qty": qty})
    report["total"] += qty

by_dept = defaultdict(list)                    # then
doc = {"departments": [{"name": d, "staff": sorted(v)} for d, v in sorted(by_dept.items())]}

Build dicts and lists as the data dictates, then serialise once. Sorting keys and lists before dumps is what makes the output stable when the source order was not.

The round trip and its limits

json.loads(json.dumps(x)) == x holds for dicts with string keys, lists, strings, ints, floats (within float precision), booleans and None. It does not hold for: tuples (become lists), non-string keys ({1: "a"} becomes {"1": "a"}), float("nan")/inf (emitted as NaN/Infinity, which is not valid JSON, unless allow_nan=False raises), and any custom object. When keys are integers, convert on the way back ({int(k): v for k, v in loaded.items()}).

Parse errors raise json.JSONDecodeError (a ValueError) with the line and column: trailing commas, single quotes and comments are all invalid JSON, however common they are in JavaScript.

Pitfalls

  • json.loads on a file object or json.load on a string (swap them).
  • Expecting tuples, sets or int keys to survive the round trip.
  • Printing a nested structure without sort_keys and comparing against expected output built in another order.
  • doc["a"]["b"] on a document that may lack "a"; chain gets or catch KeyError.
  • Mutating a list while walking it recursively.
  • Single quotes or trailing commas in hand-written JSON.

Key takeaways

  • JSON objects, arrays, strings, numbers, booleans and null map to dict, list, str, int/float, bool and None; keys are always strings.
  • loads/dumps for strings, load/dump for files; indent, sort_keys=True and ensure_ascii=False control the text.
  • Navigate optional fields with chained get and typed defaults; validate shapes with match.
  • Walk unknown depth with a recursive function dispatching on dict / list / leaf.
  • The round trip loses tuples and non-string keys; sort keys and lists for deterministic output.

Common questions

What is the difference between json.load and json.loads?

json.loads parses a string — the s stands for string — while json.load reads from a file object such as an open file or sys.stdin. Likewise, json.dumps returns a string and json.dump writes to a file.

How do I pretty-print JSON in Python?

Pass indent=2 to json.dumps for output nested by two spaces, and add sort_keys=True so that equal dicts built in different orders print identically. For any Python object, pprint.pprint does the same job without JSON's restrictions.

How do I safely access a nested dictionary key in Python?

Chain get calls with an empty default of the right type: doc.get("address", {}).get("city") returns None instead of raising KeyError at either level. When a missing field would be a bug, index directly and let the KeyError report it.

Why does json.dumps turn tuples into lists and int keys into strings?

JSON has only arrays and objects with string keys, so a tuple is written as an array and read back as a list, and {1: "a"} comes back as {"1": "a"}. Convert on the way back, for example with {int(k): v for k, v in loaded.items()}.

What does TypeError: Object of type set is not JSON serializable mean?

Sets, dates, Decimal values and class instances have no JSON form, so json.dumps raises TypeError for them. Convert the value first, such as sorted(s) for a set, or pass a default= function that turns unknown objects into something JSON can hold.

Exercises

Walk the document

Read one JSON document from standard input with json.load(sys.stdin) and print every leaf value with its path, using a recursive walk: object keys join with ., array indexes appear as [i]. Print booleans and null the Python way (True, None). Finish with the number of leaves.

Input: a JSON document. Output: <path> = <value> per leaf in document order, then leaves: <n>.

{"name": "ada", "langs": ["python", "c"], "active": true, "boss": null}

prints

name = ada
langs[0] = python
langs[1] = c
active = True
boss = None
leaves: 5

Canonical JSON

Read a JSON document and print it in canonical form — json.dumps with sorted keys and compact separators — then the number of top-level keys (0 if the document is not an object) and the maximum nesting depth, where a scalar has depth 0 and each enclosing object or array adds 1.

Input: a JSON document. Output: the canonical text, then keys: <n>, then depth: <d>.

{"b": [1, {"z": null, "a": 2}], "a": "x"}

prints

{"a":"x","b":[1,{"a":2,"z":null}]}
keys: 2
depth: 3

In this module: Dictionaries and sets

← Hashing and keys — what makes an object usable in a dict or set · Choosing a collection — the complexity table and heapq →