Python Text Tools: textwrap, difflib and Regex Lookarounds

Python's textwrap wraps text, difflib finds close matches, string.Template fills safe templates; plus regex named groups, lookarounds and sub with a function.

  • Course: Python study plan
  • Module: The standard library in depth
  • Kind: Lesson
  • Reading time: 14 min
  • Runtime: CPython 3.11

What is a lookahead in a Python regex?

A lookahead in a Python regular expression is an assertion that checks what follows the current position without consuming it. (?=...) is positive and (?!...) negative, so re.findall(r"[0-9]+(?= kg)", "5 kg, 12 lb, 7 kg") returns ['5', '7'], the numbers followed by " kg" without the unit. Lookbehinds, (?<=...) and (?<!...), check what precedes and must be fixed width.

Lesson

Beyond the string methods and the regular-expression basics of Module 5, the standard library has a second tier of text tools that turn common jobs into a call: wrapping and indenting paragraphs, computing what changed between two versions, fuzzy-matching a typed name against the valid ones, filling templates safely, and the parts of re — named groups, lookarounds, sub with a function, VERBOSE — that make a regex readable. This lesson covers textwrap, difflib, string.Template and unicodedata, then the regex features in depth.

textwrap

import textwrap

text = "The quick brown fox jumps over the lazy dog and keeps running."
textwrap.wrap(text, width=20)          # ['The quick brown fox', 'jumps over the lazy', 'dog and keeps', 'running.']
print(textwrap.fill(text, width=20))   # the same, joined with newlines
textwrap.shorten(text, width=25, placeholder="…")   # 'The quick brown fox…'
textwrap.indent("a\nb", "> ")          # '> a\n> b'
textwrap.dedent("""\
    def f():
        pass
""")                                   # removes the common leading whitespace — for code in triple-quoted strings

wrap breaks on whitespace and never splits a word unless it exceeds the width (break_long_words=False to forbid even that); initial_indent/subsequent_indent make hanging paragraphs. dedent is the fix for indented multi-line strings inside functions.

difflib

import difflib

a = "the cat sat on the mat".split()
b = "the cat lay on a mat".split()
difflib.SequenceMatcher(None, a, b).ratio()           # 0.8333… similarity in [0, 1]
for tag, i1, i2, j1, j2 in difflib.SequenceMatcher(None, a, b).get_opcodes():
    print(tag, a[i1:i2], b[j1:j2])                    # equal / replace / delete / insert with the slices
list(difflib.unified_diff(old_lines, new_lines, lineterm=""))   # a patch-style diff, like `diff -u`
difflib.get_close_matches("appel", ["apple", "ample", "apply"], n=2, cutoff=0.6)   # ['apply', 'apple'] — both score 0.8, and a tie comes out in reverse order

SequenceMatcher works on any sequences (characters, words, lines) and its ratio is the standard "how similar" number; get_close_matches is the "did you mean?" helper; unified_diff and ndiff render changes for humans. HtmlDiff produces a side-by-side table.

string.Template

from string import Template

t = Template("Hello, $name! You owe $$${amount}.")
t.substitute(name="Ada", amount="12.50")         # 'Hello, Ada! You owe $12.50.'
t.safe_substitute(name="Ada")                     # leaves $amount in place instead of raising KeyError

$identifier placeholders and $$ for a literal dollar — deliberately less powerful than f-strings and str.format, which is the point when the template comes from a user or a file: it cannot call methods or read attributes, so it cannot be abused. string also holds ascii_letters, digits, punctuation, whitespace and capwords.

unicodedata

import unicodedata

unicodedata.normalize("NFC", s)        # compose accents (compare and store in this form)
unicodedata.normalize("NFKD", "fi")     # 'fi' — compatibility decomposition
unicodedata.name("€")                  # 'EURO SIGN'
unicodedata.category("A")              # 'Lu' — letter, uppercase
"".join(c for c in unicodedata.normalize("NFKD", "café") if not unicodedata.combining(c))   # 'cafe'

The last line strips accents — the usual step before slugifying or comparing names loosely (Module 5).

re in depth

Named groups and groupdict.

m = re.search(r"(?P<key>\w+)\s*=\s*(?P<value>.*)", "port = 8080")
m["key"], m["value"], m.groupdict()          # 'port', '8080', {'key': 'port', 'value': '8080'}

Lookarounds assert without consuming:

re.findall(r"\d+(?= kg)", "5 kg, 12 lb, 7 kg")       # ['5', '7'] — followed by " kg"
re.findall(r"(?<=\$)\d+", "cost $40 or $55")          # ['40', '55'] — preceded by "$"
re.sub(r"(?<!\d)0+(?=\d)", "", "007 and 0042")        # '7 and 42' — strip leading zeros
re.findall(r"\b(?!the\b)\w+", "the cat the dog")      # words that are not "the"

(?=…) positive lookahead, (?!…) negative lookahead, (?<=…) positive lookbehind (fixed width), (?<!…) negative lookbehind. They let a pattern match at a position without including the context in the match — the tool for "a number followed by a unit, but only the number".

sub with a function transforms each match:

re.sub(r"\d+", lambda m: str(int(m.group()) * 2), "a1 b22")      # 'a2 b44'
re.sub(r"\b(\w)(\w*)", lambda m: m.group(1).upper() + m.group(2), "hello big world")   # title-case by regex

VERBOSE makes a long pattern readable:

DATE = re.compile(r"""
    (?P<year>\d{4}) -      # year
    (?P<month>\d{2}) -     # month
    (?P<day>\d{2})         # day
""", re.VERBOSE)

Whitespace and # comments are ignored in the pattern (escape a literal space as \ or [ ]).

Other tools. re.split(r"[,;]\s*", text) splits on a pattern; re.finditer yields Match objects with .start()/.end()/.span(); re.escape(s) quotes a literal for use inside a pattern; re.fullmatch for validation; re.compile(...).pattern and .flags for introspection; (?i) inline flags; non-greedy *?/+?/??; {n,m}?.

Backreferences \1 and (?P=name) match a repeat of a group: r"(\w+) \1" finds a doubled word.

Choosing between them

The string methods stay the first choice: split, partition, startswith, replace and strip are faster and clearer than any regex for a fixed delimiter or prefix. A regex earns its place when the shape matters — digits followed by a unit, a key with optional spaces around =, several alternatives at once — and re.VERBOSE with comments keeps such a pattern maintainable. difflib is for comparing and suggesting, textwrap for presenting, Template for text that came from outside; none of them is a substitute for a parser when the input has nested structure.

Pitfalls

  • textwrap.wrap on text with existing newlines — it treats them as spaces unless replace_whitespace=False.
  • Using ratio() thresholds without cutoff in get_close_matches.
  • Template.substitute raising KeyError for a missing placeholder when safe_substitute was wanted.
  • Variable-width lookbehind ((?<=\d+)) — not allowed; fixed width only.
  • re.VERBOSE swallowing a significant space in the pattern.
  • Forgetting re.escape when building a pattern from user text.

Key takeaways

  • textwrap: wrap, fill, shorten, indent, dedent; difflib: SequenceMatcher.ratio, get_opcodes, unified_diff, get_close_matches.
  • string.Template for user-supplied templates; unicodedata for normalisation and accent stripping.
  • Named groups with groupdict; lookarounds assert context without consuming; sub with a function transforms matches.
  • re.VERBOSE for readable patterns; re.escape for literals; backreferences for repeats.
  • Reach for these before writing a wrapper, a diff or a fuzzy match by hand.

Common questions

How do I find the closest matching string in Python?

Use difflib.get_close_matches(word, possibilities, n=3, cutoff=0.6), which returns up to n candidates whose similarity is at least the cutoff, best first: the "did you mean?" helper, so a typed "appel" suggests "apple". For a similarity score between two sequences, use difflib.SequenceMatcher(None, a, b).ratio(), a number from 0 to 1.

How do I remove the indentation from a multi-line string in Python?

textwrap.dedent(text) removes the whitespace common to the start of every line, the usual fix for an indented triple-quoted string inside a function. textwrap.indent(text, prefix) does the reverse, and textwrap.fill(text, width) rewraps a paragraph to a width.

What is string.Template used for in Python?

string.Template fills $name placeholders with substitute, or with safe_substitute, which leaves missing placeholders in place instead of raising KeyError. It is deliberately less powerful than f-strings and str.format: it cannot read attributes or call methods, so it is safe for templates supplied by users or files.

Can re.sub take a function as the replacement?

Yes. When the replacement is a function, re.sub calls it with each match object and inserts the string it returns, so re.sub(r"[0-9]+", lambda m: str(int(m.group()) * 2), "a1 b22") gives 'a2 b44'. Use it when the replacement depends on the matched text.

Exercises

Wrap, shorten, suggest

The first line is a width; the second a paragraph; the third a vocabulary of words; the remaining lines are words to check. Print the paragraph wrapped to the width with textwrap.fill, then shortened to the width with the placeholder […], then for each remaining word either ok if it is in the vocabulary or did you mean <best> using difflib.get_close_matches (n=1, default cutoff), or no suggestion.

Input: width, paragraph, vocabulary, words. Output: the wrapped lines, the shortened line, then one line per word.

20
The quick brown fox jumps over the lazy dog
apple banana cherry
appel
banana
zzz

prints

The quick brown fox
jumps over the lazy
dog
The quick brown […]
did you mean apple
ok
no suggestion

Regex in depth

Each line is a command. pairs <text> extracts every key=value pair with named groups and prints them as key:value space-separated in order (or none). kg <text> prints the numbers immediately followed by kg (a lookahead, so the unit is not part of the match), space-separated (or none). double <text> doubles every integer in the text with re.sub and a function.

Input: command lines. Output: one line per command.

pairs host=db port=8080 debug
kg 5 kg of rice, 12 lb of flour, 7 kg of sugar
double a1 b22 c

prints

host:db port:8080
5 7
a2 b44 c

In this module: The standard library in depth

← collections in depth — deque, Counter, OrderedDict, ChainMap and the User classes · Enums — Enum, IntEnum, StrEnum, Flag and auto →