Python Strings Explained: Immutability, Escapes, str vs repr
A Python str is an immutable sequence of Unicode code points. Literals, escapes and raw strings, indexing, comparison, str versus repr, and why += is slow.
- Course: Python study plan
- Module: Strings and text
- Kind: Lesson
- Reading time: 12 min
- Runtime: CPython 3.11
Are Python strings immutable?
Yes. A Python str is an immutable sequence of Unicode code points: once created it never changes, so s[0] = "b" raises TypeError: 'str' object does not support item assignment. Methods such as upper, replace and strip return a new string and leave the original alone. Immutability is what lets strings serve as dictionary keys and set members.
Lesson
A Python str is an immutable sequence of Unicode characters. Each of those four words matters: immutable means every "change" makes a new string; sequence means indexing, slicing, len, in and iteration all work as they do on a list; Unicode means a character is a code point, not a byte, and "é" has length 1; and string means the type comes with the richest method set in the language. This lesson covers literals and escapes, the sequence operations, comparison, str versus repr, and the concatenation cost that decides how output should be built.
Literals
a = 'single'
b = "double" # same thing; pick one and be consistent
c = "it's" # the other quote avoids escaping
d = 'say "hi"'
e = "line one\nline two\ttabbed" # escapes: \n newline, \t tab, \\ backslash, \" quote
f = r"C:\new\folder" # raw: backslashes are literal (regexes, Windows paths)
g = """A multi-line
string keeps its newlines."""
h = "adjacent " "literals " "join" # implicit concatenation at compile time
Escapes are interpreted in ordinary literals: \n, \t, \\, \', \", \x41 (hex byte), \u00e9 (Unicode code point), \N{EURO SIGN}. A raw string turns them off — except that a raw string still cannot end in an odd backslash. Triple-quoted strings span lines and are the docstring syntax. Adjacent literals are joined by the compiler, which is how a long message is split across lines inside parentheses without +.
Sequence operations
s = "python"
len(s) # 6
s[0], s[-1] # 'p', 'n' — negative indexes count from the end
s[2:4] # 'th' — slicing, half-open (next lesson)
"th" in s # True — substring test
s + "3" # 'python3'
"-" * 10 # '----------'
for ch in s: # iterates characters
...
list(s) # ['p', 'y', 't', 'h', 'o', 'n']
Indexing past the end raises IndexError; slicing past the end does not (it clips). There is no separate character type — s[0] is a string of length 1. in tests for a substring, not just a character: "tho" in "python" is True.
Immutability
s = "cat"
s[0] = "b" # TypeError: 'str' object does not support item assignment
s = "b" + s[1:] # 'bat' — a new string bound to the same name
Every method that appears to modify — upper, replace, strip — returns a new string and leaves the original alone; the classic bug is calling one and discarding the result (s.upper() on its own line does nothing). Immutability is what makes strings usable as dictionary keys and set members, and it is why building a long string by repeated += is expensive: each step copies the whole accumulated string. Collect the pieces in a list and "".join them once (Module 3 lesson 5).
Comparison
Strings compare by code point, left to right, first difference decides; a prefix sorts before the longer string:
"apple" < "banana" # True
"Zebra" < "apple" # True — uppercase letters have smaller code points than lowercase
"abc" < "abd" # True
"ab" < "abc" # True
"10" < "9" # True — text, not numbers
Case-insensitive comparison is a.lower() == b.lower() — or, correctly for every language, a.casefold() == b.casefold(). Comparing numeric strings as numbers means converting them first: sorted(nums, key=int). == compares content, never identity; two equal strings may or may not be the same object, and is is not for strings.
str and repr
str(x) is the readable form, what print shows; repr(x) is the unambiguous form, what the REPL shows and what you want in a debug message:
s = "a\tb\n"
print(str(s)) # a b (a tab and a newline, invisible)
print(repr(s)) # 'a\tb\n'
print(f"{s!r}") # the same, inside an f-string
repr of a string is a valid literal that would recreate it, quotes and escapes included. When a program's output is wrong by "nothing visible", repr shows the trailing newline or the tab.
The methods, first pass
The next lesson goes through them; the ones every program needs are strip() (remove surrounding whitespace), split() (to a list of words), join(iterable) (from a list back to a string), upper()/lower(), replace(old, new), startswith/endswith, find, and count. Note the direction of join: it is a method of the separator — ", ".join(names) — because the separator is the one string that is always a string, while the pieces might be any iterable.
Pitfalls
s.strip()(or any method) called without keeping the result.s[0] = "x". Build a new string."10" < "9"andsorted(["10", "9"])— text order, not numeric.+=in a loop over many pieces; use a list andjoin.- Treating
s[i]as a char type: it is astr, ands[i] == "a"is the comparison. - Confusing
strandreprwhen printing a debug value.
Key takeaways
- Strings are immutable sequences of code points: index, slice,
len,in, iterate; every change is a new string. - Literals: single or double quotes, escapes, raw strings for backslashes, triple quotes for multi-line, adjacent literals join.
- Comparison is by code point (uppercase before lowercase, text order for digits);
casefoldfor case-insensitive. stris readable,repris unambiguous — usereprwhen debugging whitespace.sep.join(pieces)builds a string; repeated+=copies every time.
Common questions
How do I change a character in a Python string?
Build a new string, since a string cannot be changed in place: s = "b" + s[1:] turns "cat" into "bat", and s.replace(old, new) returns a copy with the substitution made. Calling a method without keeping its result, like s.upper() on its own line, changes nothing.
What is the difference between str and repr in Python?
str(x) is the readable form that print shows; repr(x) is the unambiguous form the REPL shows. For a string, repr adds quotes and writes tabs and newlines as escapes, so whitespace you cannot see becomes visible. In an f-string, !r applies it: f"{s!r}".
What is a raw string in Python?
A raw string, written with an r prefix, treats backslashes as ordinary characters instead of starting escapes such as newline or tab. It is the usual form for regular expressions and Windows paths. One limit remains: a raw string cannot end in an odd number of backslashes.
How are strings compared in Python?
Character by character by Unicode code point, and the first difference decides; a prefix sorts before the longer string. Uppercase letters come before lowercase, so "Zebra" < "apple" is True, and digits compare as text, so "10" < "9" is True. Use casefold() on both sides to ignore case.
How do I write a multi-line string in Python?
Use triple quotes, """ or ''', which keep the newlines typed between them; they are also the docstring syntax. To split one long line of text across several source lines instead, put adjacent string literals inside parentheses, and the compiler joins them into one string.
Exercises
Palindromes, normalised
Read lines until the end of input and decide whether each is a palindrome when only letters and digits are considered and case is ignored: build the normalised string with casefold() and a filter on isalnum(), then compare it with its reverse [::-1].
Input: lines of text. Output: yes or no per line.
A man, a plan, a canal: Panama
hello
prints
yes
noText order versus numeric order
Strings compare by code point, which is not numeric order and puts uppercase before lowercase. Read a line of tokens and print them sorted as text; then, if every token parses as an integer, print them sorted numerically as well (using key=int), otherwise print numeric: n/a.
Input: one line of tokens. Output: text: <tokens> then numeric: <tokens> or numeric: n/a.
10 9 100
prints
text: 10 100 9
numeric: 9 10 100In this module: Strings and text
- String basics — an immutable sequence of characters (this lesson)
- Slicing and the string methods
- Formatting — f-strings and the format-spec mini-language
- Characters, code points, bytes and Unicode
- Parsing input — from lines and tokens to values
- Regular expressions — the re module
- Checkpoint — Strings and text