Python Strings Quiz: Slicing, F-Strings, Unicode and Regex

Test your Python string skills with 12 questions and three programs on slicing, string methods, f-strings, Unicode and bytes, parsing and regular expressions.

  • Course: Python study plan
  • Module: Strings and text
  • Kind: Checkpoint — cleared at 70%
  • Reading time: 25 min
  • Runtime: CPython 3.11

Checkpoint — Strings and text is the checkpoint that closes the Strings and text module: a graded quiz and whole-program exercises, passed at 70%.

Instructions

This checkpoint covers the whole module: strings as immutable sequences of code points, slicing and the method set, f-strings and the format-spec mini-language, characters and the text/bytes boundary with the three digit tests, the parsing shapes for tokens, fields, key=value pairs, records and blocks, and regular expressions with groups, substitution and the greedy rule.

How it works. Twelve questions and three programs. You need 70% on the questions and every program accepted to clear the module. You can retake it as often as you like; your best score counts.

Before you start, make sure you can answer these from memory:

  • What is "abcdef"[1:5:2], and what does [::-1] do?
  • Why is "file.txt".strip(".txt") wrong, and what is right?
  • What do f"{42:06d}", f"{3.14159:.2f}" and f"{'ab':>5}" produce?
  • What is the difference between len("é") and len("é".encode())?
  • Which of isdecimal, isdigit, isnumeric accepts "²", and which accepts "-1"?
  • What does partition("=") return when the separator is absent?
  • What is the difference between re.match and re.search, and what do they return when nothing matches?

The three programs are a text-statistics report built from the method set and formatted as a table, a key=value configuration parser with typed values and clear rejections, and a log extractor that pulls timestamps, levels and messages out of free text with named groups.

Common questions

What is "abcdef"[1:5:2] in Python?

It is "bd": the slice starts at index 1, stops before index 5 and takes every second character, so it picks indexes 1 and 3. A step of -1, as in [::-1], takes every character from the end and reverses the string.

Why is "file.txt".strip(".txt") wrong?

strip treats its argument as a set of characters to remove from both ends, not as a suffix, so it also strips any t, x or . at the edges; here it returns "file" only by luck. Use "file.txt".removesuffix(".txt"), available since Python 3.9.

What do f"{42:06d}", f"{3.14159:.2f}" and f"{'ab':>5}" produce?

000042, 3.14 and ab preceded by three spaces: an integer zero-padded to six characters, a float rounded to two decimals, and a string right-aligned in a field five characters wide.

What is the difference between len("é") and len("é".encode())?

len("é") is 1, because a str counts Unicode code points; len("é".encode()) is 2, because UTF-8 encodes é in two bytes. A string's length is never its size in bytes.

Exercises

Text statistics

Read all of standard input and report, as a two-column table with labels left-aligned in 14 characters and values right-aligned in 8: lines (from splitlines), words (whitespace tokens), chars (every character), unique (distinct words after casefold and stripping punctuation from both ends; empty results do not count), longest (the first longest cleaned word, left-aligned like the labels' column but printed in the value column right-aligned), and avg length (mean length of cleaned words, two decimals, 0.00 for none).

Input: any text. Output: six table rows.

The cat sat. The cat ran!

prints

lines                1
words                6
chars               26
unique               4
longest            the
avg length        3.00

Query strings

Parse query strings, one per line: fields separated by &, each key=value or a bare key (a flag). A + in a value stands for a space. Type each value as in the config parser (int, float, bool for true/false, else str); a value containing commas is a list of strings printed joined with |; a bare key is bool True. A key seen twice on the same line prints duplicate: <key> instead of a second entry. Print a blank-free block per line: the entries in order, then --.

Input: lines. Output: per line, the entries then --.

name=ada+lovelace&age=36&tags=a,b,c&admin&age=1

prints

name: str ada lovelace
age: int 36
tags: list a|b|c
admin: bool True
duplicate: age
--

Access-log summary

Each line of an access log is <ip> - - [<timestamp>] "<method> <path> HTTP/1.1" <status> <bytes>. With one regular expression using named groups, parse every line (skip lines that do not match) and print: the number of requests per status code in ascending numeric order, the total bytes, and the most requested path (ties broken alphabetically).

Input: lines. Output: status <code>: <count> lines, then bytes: <total>, then top: <path> (or top: none).

10.0.0.1 - - [01/May/2024:12:00:00 +0000] "GET /index.html HTTP/1.1" 200 512
10.0.0.2 - - [01/May/2024:12:00:01 +0000] "GET /missing HTTP/1.1" 404 0
10.0.0.1 - - [01/May/2024:12:00:02 +0000] "POST /index.html HTTP/1.1" 200 128
not a log line

prints

status 200: 2
status 404: 1
bytes: 640
top: /index.html

In this module: Strings and text

← Regular expressions — the re module · Lists — the mutable sequence →