How to Parse Input in Python: Split, Partition and Validate
Parse Python input with split, strip and partition, then convert with int or float. Templates for tokens, fields, key=value pairs, records and bulk reads.
- Course: Python study plan
- Module: Strings and text
- Kind: Lesson
- Reading time: 14 min
- Runtime: CPython 3.11
How do you parse input in Python?
Parse Python input by splitting the text into pieces and converting each one strictly: split() for whitespace-separated tokens, split(",") plus strip() for delimited fields, partition("=") for key=value pairs, then int() or float() with ValueError caught for bad values. For large inputs, sys.stdin.read().split() reads every token in one call.
Lesson
Input arrives as text and the program needs values, and everything between is parsing. Most of it is three methods — split, strip, partition — plus int and float with ValueError caught, arranged into a handful of shapes: tokens on a line, delimited fields, key=value pairs, a count followed by records, blocks separated by blank lines, and the occasional format that needs a small hand-written scanner. This lesson gives each shape a template, states the validation rule (convert strictly, reject clearly), and shows where the bulk-read idiom sys.stdin.read().split() beats reading line by line.
Tokens on a line
nums = list(map(int, input().split()))
a, b = map(int, input().split()) # exactly two — ValueError otherwise
first, *rest = input().split() # at least one
split() without an argument handles any amount of whitespace and drops the ends. The unpacking forms are also validation: the wrong number of tokens raises ValueError: not enough values to unpack or too many values, which is the right failure for malformed input.
Delimited fields
line = "ada, 36, london"
name, age, city = [f.strip() for f in line.split(",")]
split(",") keeps empty fields, so "a,,c" has three; strip each field, because ", " after a comma is common. When fields can contain the delimiter (quoted CSV), do not split by hand — the csv module (Module 14) handles quoting. A fixed number of fields can be unpacked; a variable number stays a list.
key=value pairs
settings = {}
for tok in "host=db port=5432 debug".split():
key, sep, value = tok.partition("=")
settings[key] = value if sep else True
partition splits at the first separator and returns three parts, with an empty separator when it is absent — so a value containing = survives ("a=b=c" → ("a", "=", "b=c")), and a bare flag is detected by sep being empty. split("=", 1) is the two-element alternative but raises on unpacking when the separator is missing.
A count, then records
n = int(input())
rows = []
for _ in range(n):
name, qty, price = input().split()
rows.append((name, int(qty), float(price)))
Convert as you read so that a bad line fails at the line, not later in the middle of a calculation. When the count line is absent and the records run to the end of input, for line in sys.stdin replaces the range.
Blocks separated by blank lines
import sys
text = sys.stdin.read()
blocks = [b.splitlines() for b in text.strip().split("\n\n") if b.strip()]
Read everything, split on the double newline, then split each block into lines. strip() on the whole text first avoids a leading or trailing empty block. A variant, one record per block with key: value lines, is dict(line.split(": ", 1) for line in block).
Bulk tokens
import sys
data = sys.stdin.read().split()
it = iter(data)
n = int(next(it))
values = [int(next(it)) for _ in range(n)]
When the input is large — tens of thousands of numbers — reading it in one call and walking an iterator is several times faster than input() per line, and it does not care how the numbers are wrapped across lines. next(it) raises StopIteration if the data runs out, which a wrapper can turn into a clear message. Module 20 uses this as the interview template.
Validating
Convert strictly and catch the failure at the point of conversion:
def parse_age(tok):
try:
age = int(tok)
except ValueError:
raise ValueError(f"age must be an integer, got {tok!r}") from None
if not 0 <= age <= 150:
raise ValueError(f"age out of range: {age}")
return age
Two rules. Do not pre-check with isdigit — it rejects -5 and accepts ²; int() is the definition of "parses as an integer". And do not let a bad value through as 0 or None silently — either raise, or record the failure and report it (bad lines: 3), depending on what the program is for. In a judged exercise the prompt says which; "print invalid and continue" is the common contract.
A hand-written scanner
Some formats are neither tokens nor fields: a number followed directly by a unit (12kg), an expression without spaces (3+4*2), a run-length code (a3b2). The tool is an index walked by hand, reading a run of characters of one class at a time:
def scan_number_and_unit(s):
i = 0
while i < len(s) and (s[i].isdigit() or s[i] == "."):
i += 1
return float(s[:i]), s[i:].strip()
def run_length_decode(code):
out, i = [], 0
while i < len(code):
ch = code[i]
i += 1
j = i
while j < len(code) and code[j].isdigit():
j += 1
out.append(ch * int(code[i:j]))
i = j
return "".join(out)
The shape is always: a position i; a loop that advances i past one lexical unit; the slice s[start:i] as the token. Regular expressions (next lesson) express the same thing declaratively and are the better tool once the format has more than two token kinds.
Pitfalls
split(" ")for whitespace tokens — it keeps empties between double spaces.- Forgetting to
strip()fields aftersplit(","). int(line)wherelinecame fromsys.stdiniteration and still has\n— actually fine (inttolerates whitespace), butline == "END"is not; compare againstline.rstrip("\n").- Reading
nand then readingntokens on one line withinput()per token —input()reads a line. - Checking
isdigit()beforeint(). - Silent defaults for unparsable values.
Key takeaways
split()for whitespace tokens,split(sep)for fields (strip each),partitionforkey=value.- Unpacking validates token counts;
int/floatvalidate values — catchValueError, never pre-check withisdigit. - Count-then-records with
range(n); records to the end withfor line in sys.stdin; blocks withread().split("\n\n"). sys.stdin.read().split()plus an iterator is the fast bulk reader.- A hand scanner is an index advanced past one token class at a time; regexes take over when the grammar grows.
Common questions
How do I split a key=value string in Python?
Use partition: key, sep, value = tok.partition("=") splits at the first = and always returns three parts, so a value containing = survives and a missing separator shows up as an empty sep. split("=", 1) also works, but unpacking its result fails when the = is absent.
What does ValueError: not enough values to unpack mean?
The names on the left of an unpacking assignment outnumbered the items on the right; with input, the line had fewer tokens than expected, as in a, b = input().split() on a one-word line. too many values to unpack is the opposite case. Both are the right failure for malformed input.
How do I read input fast in Python?
Read everything at once with sys.stdin.read().split() and walk the tokens with an iterator: it = iter(data), then int(next(it)) for each value. For tens of thousands of numbers this is several times faster than calling input() per line, and it does not care how the values are wrapped across lines.
Should I check isdigit() before calling int() in Python?
No. isdigit() rejects valid integers such as "-5" and accepts characters that int() refuses, such as "²". int() itself defines what parses as an integer, so call it inside try and turn the ValueError into a clear message.
How do I parse a comma-separated line in Python?
[f.strip() for f in line.split(",")] splits at each comma, keeps empty fields, and strips the spaces that often follow a comma. When fields can contain the delimiter inside quotes, use the csv module rather than splitting by hand.
Exercises
Config parser
Parse a configuration read until the end of input. Blank lines and lines starting with # are ignored. Every other line must be key = value (spaces around = optional) — split with partition, and print invalid: <line> when there is no =. Type each value: an integer, else a float, else true/false (any case) as a boolean, else a string. Print the entries in input order.
Input: lines. Output: <key>: <type> <value> per entry (types int, float, bool, str; booleans print as True/False).
# server
host = localhost
port=8080
ratio = 0.75
debug = true
name
prints
host: str localhost
port: int 8080
ratio: float 0.75
debug: bool True
invalid: nameRun-length codec
Each line is E <text> (encode) or D <code> (decode). Encoding turns runs of a repeated character into the character followed by its count (aaabcc → a3b1c2). Decoding reverses it; counts may have several digits (x10). Write the decoder as a hand scanner: an index that reads one character, then the run of digits after it. Text after the command is taken literally (it may contain spaces).
Input: lines. Output: one line per input line (an empty result prints (empty)).
E aaabcc
D a3b1c2
prints
a3b1c2
aaabccIn this module: Strings and text
- String basics — an immutable sequence of characters
- Slicing and the string methods
- Formatting — f-strings and the format-spec mini-language
- Characters, code points, bytes and Unicode
- Parsing input — from lines and tokens to values (this lesson)
- Regular expressions — the re module
- Checkpoint — Strings and text
← Characters, code points, bytes and Unicode · Regular expressions — the re module →