Python Regex Explained: re.search, re.match, Groups and sub
The Python re module finds, extracts, replaces and splits text by pattern. search vs match, findall, named groups, sub, flags and greedy vs lazy quantifiers.
- Course: Python study plan
- Module: Strings and text
- Kind: Lesson
- Reading time: 15 min
- Runtime: CPython 3.11
How do you use regular expressions in Python?
Python's regular expressions live in the re module: re.search(pattern, text) finds the first match anywhere and returns a Match object or None, re.findall returns every match as a list, re.sub replaces matches and re.split splits on them. Write patterns as raw strings with an r prefix so that their backslashes reach the regex engine intact.
Lesson
A regular expression describes a set of strings by a pattern — "one or more digits", "a word, a colon, then anything" — and the re module finds, extracts, replaces and splits by such patterns. It is the right tool when the format is more than one delimiter deep, when you need "the numbers anywhere in this text", or when validation is a shape rather than a value. It is the wrong tool for anything split or startswith can do, and for parsing nested structures. This lesson gives the syntax that covers most uses, the six functions, groups, flags, and the greedy-versus-lazy rule.
The functions
import re
text = "order 66 shipped 2024-05-01 to ada"
re.search(r"\d+", text) # first match anywhere: <re.Match ... match='66'>
re.match(r"\w+", text) # match only at the *start*: 'order'
re.fullmatch(r"\w+", "order") # the whole string must match
re.findall(r"\d+", text) # every match as a list: ['66', '2024', '05', '01']
re.finditer(r"\d+", text) # every match as Match objects, lazily
re.sub(r"\d", "#", text) # replace: 'order ## shipped ####-##-## to ada'
re.split(r"[\s-]+", text) # split on a pattern
search, match and fullmatch return a Match object or None — always test before using it. m.group() is the matched text, m.start()/m.end() its position, m.group(1) the first parenthesised group. Patterns are written as raw strings (r"...") so that backslashes reach the regex engine intact.
The syntax
| Pattern | Matches | |
|---|---|---|
. | any character except newline | |
\d \w \s | digit, word character (letters, digits, _), whitespace | |
\D \W \S | the complements | |
[abc] [a-z] [^0-9] | a set, a range, a negated set | |
^ $ | start and end of the string (or line, with re.M) | |
\b | a word boundary | |
x* x+ x? | zero or more, one or more, zero or one | |
x{3} x{2,5} x{2,} | exactly, between, at least | |
| `a\ | b` | alternation |
(...) | a group: captures, and groups a quantifier | |
(?:...) | a non-capturing group | |
(?P<name>...) | a named group | |
(?=...) (?!...) | lookahead: followed by / not followed by, without consuming | |
\. \( \\ | a literal special character |
re.fullmatch(r"[A-Z]{3}-\d{4}", "ABC-1234") # a plate
re.search(r"\b(\w+)\s+\1\b", "the the cat") # a repeated word (\1 = group 1 again)
re.findall(r"[\w.]+@[\w.]+\.\w+", text) # email-shaped things (not validation)
Groups
m = re.search(r"(\d{4})-(\d{2})-(\d{2})", text)
m.group(0) # '2024-05-01' — the whole match
m.group(1), m.group(3) # '2024', '01'
m.groups() # ('2024', '05', '01')
m = re.search(r"(?P<year>\d{4})-(?P<month>\d{2})", text)
m.group("year") # '2024'
m.groupdict() # {'year': '2024', 'month': '05'}
findall returns the groups as tuples when the pattern has groups (and just group 1 when it has one), which surprises people who wanted the whole match — use (?:...) for grouping without capturing, or finditer and m.group().
Substitution
re.sub(r"(\w+)@(\w+)", r"\2 at \1", "ada@example") # 'example at ada' — backreferences
re.sub(r"\d+", lambda m: str(int(m.group()) * 2), "a1 b22") # 'a2 b44' — a function per match
re.sub(r"\s+", " ", " many spaces ").strip() # 'many spaces'
The replacement string understands \1 and \g<name>; a function receives the Match and returns the replacement, which is how a transformation (upper-casing, arithmetic) is applied to each match.
Greedy and lazy
Quantifiers are greedy: they match as much as possible and back off only if the rest of the pattern fails. <.*> on <a><b> matches the whole string. Add ? to make a quantifier lazy — <.*?> matches <a> — or, usually better, forbid the closing character in the set: <[^>]*>. The negated set is faster and clearer than laziness.
Flags and compiling
re.findall(r"^\w+", text, flags=re.M) # ^ and $ per line
re.search(r"hello", "HELLO", re.I) # ignore case
re.compile(r"""
(?P<key>\w+) # the key
\s*=\s*
(?P<val>.*) # the value
""", re.X) # verbose: whitespace and comments ignored
re.compile(pattern) returns a pattern object with the same methods (p.search(s)); the module functions cache recent patterns, so compiling is about readability and reuse, not speed, except in a tight loop. re.S (dotall) lets . match newlines.
When not to use a regex
s.startswith("x"), "," in s, s.split(","), s.isdigit() and s.replace(a, b) are each clearer and faster than the equivalent regex. Nested or balanced structure (matching brackets, HTML) is beyond regular languages; use a parser. Validation of real-world formats (emails, URLs) with a regex is a rough filter, never a proof. And a regex that took ten minutes to write will take the next reader ten minutes to read — the re.X flag with comments is the mitigation.
Pitfalls
- Forgetting the
rprefix and losing\b(backspace) or\1. - Using the result of
search/matchwithout checking forNone. matchwhensearchwas meant —matchanchors at the start.findallwith groups returning tuples instead of whole matches.- A greedy
.*spanning far more than intended. - Special characters unescaped:
.matches anything,+quantifies;re.escape(s)escapes a literal.
Key takeaways
searchfinds anywhere,matchat the start,fullmatchthe whole; all returnMatchorNone— test it.findall/finditerfor all matches,sub(with backreferences or a function) to replace,spliton a pattern.- Groups capture with
(...), name with(?P<n>...), skip capturing with(?:...);\1refers back. - Quantifiers are greedy;
?makes them lazy; a negated set[^>]*is usually better than either. - Raw strings always; flags
I,M,S,X; prefer a string method when one does the job.
Common questions
What is the difference between re.match and re.search in Python?
re.match only matches at the start of the string, while re.search finds the first match anywhere in it; re.fullmatch requires the whole string to match. All three return a Match object, or None when nothing matches, so test the result before calling .group() on it.
How do groups work in Python regular expressions?
Parentheses capture: after a match, m.group(1) is the first group's text, m.group(0) the whole match and m.groups() a tuple of every group. (?P<year>...) names a group, read back with m.group("year") or m.groupdict(), and (?:...) groups without capturing.
Why does re.findall return tuples?
When a pattern contains capturing groups, findall returns the groups instead of the whole match: a list of tuples for several groups, or group 1's text alone for one. Make grouping parentheses non-capturing with (?:...), or use re.finditer and m.group() to get whole matches.
What is the difference between greedy and lazy matching in regex?
Quantifiers such as * and + are greedy: they match as much as possible, so <.*> on <a><b> matches the whole string. Adding ? makes one lazy, and <.*?> matches only <a>, but forbidding the closing character, as in <[^>]*>, is usually faster and clearer.
Why should regex patterns be raw strings in Python?
Without the r prefix, Python processes backslash escapes before the regex engine sees the pattern, so a word boundary turns into a backspace character and a backreference to group 1 into a control character. A raw string passes every backslash through unchanged.
Exercises
Extract the numbers
For each line of text, find every number — an optional minus sign, digits, and optionally a decimal point with more digits — with one re.findall pattern that uses a non-capturing group for the fraction. Print the matched texts and their sum (as floats, two decimals).
Input: lines of text. Output: found: <matches> (or found: none) then sum: <total>, per line.
Order 66 costs 12.50, tax 2.5
no digits here
prints
found: 66 12.50 2.5
sum: 81.00
found: none
sum: 0.00Log lines with named groups
A log line looks like 2024-05-01 12:30:45 [ERROR] Disk full on /dev/sda1. Write one pattern with named groups date, time, level and msg, apply it with fullmatch, and print <level> @ <time>: <msg> for matching lines and skip for the others. After all lines, print the count per level in alphabetical order as levels: ERROR=1 INFO=2 (or levels: none).
Input: lines. Output: one line per input line, then the summary.
2024-05-01 12:30:45 [ERROR] Disk full on /dev/sda1
2024-05-01 12:31:00 [INFO] Retrying
garbage
prints
ERROR @ 12:30:45: Disk full on /dev/sda1
INFO @ 12:31:00: Retrying
skip
levels: ERROR=1 INFO=1In this module: Strings and text
- String basics — an immutable sequence of characters
- Slicing and the string methods
- Formatting — f-strings and the format-spec mini-language
- Characters, code points, bytes and Unicode
- Parsing input — from lines and tokens to values
- Regular expressions — the re module (this lesson)
- Checkpoint — Strings and text
← Parsing input — from lines and tokens to values · Checkpoint — Strings and text →