How to Parse Input in Python: Split, Partition and Validate

Parse Python input with split, strip and partition, then convert with int or float. Templates for tokens, fields, key=value pairs, records and bulk reads.

  • Course: Python study plan
  • Module: Strings and text
  • Kind: Lesson
  • Reading time: 14 min
  • Runtime: CPython 3.11

How do you parse input in Python?

Parse Python input by splitting the text into pieces and converting each one strictly: split() for whitespace-separated tokens, split(",") plus strip() for delimited fields, partition("=") for key=value pairs, then int() or float() with ValueError caught for bad values. For large inputs, sys.stdin.read().split() reads every token in one call.

Lesson

Input arrives as text and the program needs values, and everything between is parsing. Most of it is three methods — split, strip, partition — plus int and float with ValueError caught, arranged into a handful of shapes: tokens on a line, delimited fields, key=value pairs, a count followed by records, blocks separated by blank lines, and the occasional format that needs a small hand-written scanner. This lesson gives each shape a template, states the validation rule (convert strictly, reject clearly), and shows where the bulk-read idiom sys.stdin.read().split() beats reading line by line.

Tokens on a line

nums = list(map(int, input().split()))
a, b = map(int, input().split())        # exactly two — ValueError otherwise
first, *rest = input().split()          # at least one

split() without an argument handles any amount of whitespace and drops the ends. The unpacking forms are also validation: the wrong number of tokens raises ValueError: not enough values to unpack or too many values, which is the right failure for malformed input.

Delimited fields

line = "ada, 36, london"
name, age, city = [f.strip() for f in line.split(",")]

split(",") keeps empty fields, so "a,,c" has three; strip each field, because ", " after a comma is common. When fields can contain the delimiter (quoted CSV), do not split by hand — the csv module (Module 14) handles quoting. A fixed number of fields can be unpacked; a variable number stays a list.

key=value pairs

settings = {}
for tok in "host=db port=5432 debug".split():
    key, sep, value = tok.partition("=")
    settings[key] = value if sep else True

partition splits at the first separator and returns three parts, with an empty separator when it is absent — so a value containing = survives ("a=b=c" → ("a", "=", "b=c")), and a bare flag is detected by sep being empty. split("=", 1) is the two-element alternative but raises on unpacking when the separator is missing.

A count, then records

n = int(input())
rows = []
for _ in range(n):
    name, qty, price = input().split()
    rows.append((name, int(qty), float(price)))

Convert as you read so that a bad line fails at the line, not later in the middle of a calculation. When the count line is absent and the records run to the end of input, for line in sys.stdin replaces the range.

Blocks separated by blank lines

import sys
text = sys.stdin.read()
blocks = [b.splitlines() for b in text.strip().split("\n\n") if b.strip()]

Read everything, split on the double newline, then split each block into lines. strip() on the whole text first avoids a leading or trailing empty block. A variant, one record per block with key: value lines, is dict(line.split(": ", 1) for line in block).

Bulk tokens

import sys
data = sys.stdin.read().split()
it = iter(data)
n = int(next(it))
values = [int(next(it)) for _ in range(n)]

When the input is large — tens of thousands of numbers — reading it in one call and walking an iterator is several times faster than input() per line, and it does not care how the numbers are wrapped across lines. next(it) raises StopIteration if the data runs out, which a wrapper can turn into a clear message. Module 20 uses this as the interview template.

Validating

Convert strictly and catch the failure at the point of conversion:

def parse_age(tok):
    try:
        age = int(tok)
    except ValueError:
        raise ValueError(f"age must be an integer, got {tok!r}") from None
    if not 0 <= age <= 150:
        raise ValueError(f"age out of range: {age}")
    return age

Two rules. Do not pre-check with isdigit — it rejects -5 and accepts ²; int() is the definition of "parses as an integer". And do not let a bad value through as 0 or None silently — either raise, or record the failure and report it (bad lines: 3), depending on what the program is for. In a judged exercise the prompt says which; "print invalid and continue" is the common contract.

A hand-written scanner

Some formats are neither tokens nor fields: a number followed directly by a unit (12kg), an expression without spaces (3+4*2), a run-length code (a3b2). The tool is an index walked by hand, reading a run of characters of one class at a time:

def scan_number_and_unit(s):
    i = 0
    while i < len(s) and (s[i].isdigit() or s[i] == "."):
        i += 1
    return float(s[:i]), s[i:].strip()

def run_length_decode(code):
    out, i = [], 0
    while i < len(code):
        ch = code[i]
        i += 1
        j = i
        while j < len(code) and code[j].isdigit():
            j += 1
        out.append(ch * int(code[i:j]))
        i = j
    return "".join(out)

The shape is always: a position i; a loop that advances i past one lexical unit; the slice s[start:i] as the token. Regular expressions (next lesson) express the same thing declaratively and are the better tool once the format has more than two token kinds.

Pitfalls

  • split(" ") for whitespace tokens — it keeps empties between double spaces.
  • Forgetting to strip() fields after split(",").
  • int(line) where line came from sys.stdin iteration and still has \n — actually fine (int tolerates whitespace), but line == "END" is not; compare against line.rstrip("\n").
  • Reading n and then reading n tokens on one line with input() per token — input() reads a line.
  • Checking isdigit() before int().
  • Silent defaults for unparsable values.

Key takeaways

  • split() for whitespace tokens, split(sep) for fields (strip each), partition for key=value.
  • Unpacking validates token counts; int/float validate values — catch ValueError, never pre-check with isdigit.
  • Count-then-records with range(n); records to the end with for line in sys.stdin; blocks with read().split("\n\n").
  • sys.stdin.read().split() plus an iterator is the fast bulk reader.
  • A hand scanner is an index advanced past one token class at a time; regexes take over when the grammar grows.

Common questions

How do I split a key=value string in Python?

Use partition: key, sep, value = tok.partition("=") splits at the first = and always returns three parts, so a value containing = survives and a missing separator shows up as an empty sep. split("=", 1) also works, but unpacking its result fails when the = is absent.

What does ValueError: not enough values to unpack mean?

The names on the left of an unpacking assignment outnumbered the items on the right; with input, the line had fewer tokens than expected, as in a, b = input().split() on a one-word line. too many values to unpack is the opposite case. Both are the right failure for malformed input.

How do I read input fast in Python?

Read everything at once with sys.stdin.read().split() and walk the tokens with an iterator: it = iter(data), then int(next(it)) for each value. For tens of thousands of numbers this is several times faster than calling input() per line, and it does not care how the values are wrapped across lines.

Should I check isdigit() before calling int() in Python?

No. isdigit() rejects valid integers such as "-5" and accepts characters that int() refuses, such as "²". int() itself defines what parses as an integer, so call it inside try and turn the ValueError into a clear message.

How do I parse a comma-separated line in Python?

[f.strip() for f in line.split(",")] splits at each comma, keeps empty fields, and strips the spaces that often follow a comma. When fields can contain the delimiter inside quotes, use the csv module rather than splitting by hand.

Exercises

Config parser

Parse a configuration read until the end of input. Blank lines and lines starting with # are ignored. Every other line must be key = value (spaces around = optional) — split with partition, and print invalid: <line> when there is no =. Type each value: an integer, else a float, else true/false (any case) as a boolean, else a string. Print the entries in input order.

Input: lines. Output: <key>: <type> <value> per entry (types int, float, bool, str; booleans print as True/False).

# server
host = localhost
port=8080
ratio = 0.75
debug = true
name

prints

host: str localhost
port: int 8080
ratio: float 0.75
debug: bool True
invalid: name

Run-length codec

Each line is E <text> (encode) or D <code> (decode). Encoding turns runs of a repeated character into the character followed by its count (aaabcc → a3b1c2). Decoding reverses it; counts may have several digits (x10). Write the decoder as a hand scanner: an index that reads one character, then the run of digits after it. Text after the command is taken literally (it may contain spaces).

Input: lines. Output: one line per input line (an empty result prints (empty)).

E aaabcc
D a3b1c2

prints

a3b1c2
aaabcc

In this module: Strings and text

← Characters, code points, bytes and Unicode · Regular expressions — the re module →