Python Files Quiz: CSV, JSON, pathlib and sqlite3 Practice
Test yourself: 12 questions and three programs on Python files, encodings, pathlib, CSV, custom JSON encoders, struct, hashlib and sqlite3.
- Course: Python study plan
- Module: Files and data formats
- Kind: Checkpoint — cleared at 70%
- Reading time: 25 min
- Runtime: CPython 3.11
Checkpoint — Files and data formats is the checkpoint that closes the Files and data formats module: a graded quiz and whole-program exercises, passed at 70%.
Instructions
This checkpoint covers the whole module: open with modes, encodings and with, line iteration and buffered writes, pathlib for building, querying and listing paths, CSV with DictReader/DictWriter and the quoting rules, JSON with custom encoders, decoders and dataclass round trips, bytes with struct, byte order, base64 and hashlib, and sqlite3 with parameters, transactions and aggregation.
How it works. Twelve questions and three programs. You need 70% on the questions and every program accepted to clear the module. You can retake it as often as you like; your best score counts.
Before you start, make sure you can answer these from memory:
- Why state
encoding="utf-8"on every textopen, and what does"w"do to an existing file? - What is
Path("a") / "b.tar.gz".suffix, and how do you get both suffixes? - Why does
csvneednewline="", and what type is every field it reads? - How do you serialise a
dateor asetwithjson.dumps? - What does
struct.pack("<I", 1)produce, and what would">I"give? - Why must SQL never be built with an f-string, and what does
with con:guarantee? - Why add
ORDER BYto a query whose output is printed?
The three programs are a CSV-to-JSON converter that writes a file and reads it back, a binary record store written with struct and verified with a SHA-256 digest, and an in-memory sqlite3 report built from lines of input with a grouped, ordered query.
Common questions
What does struct.pack("<I", 1) produce, and what would ">I" give?
struct.pack("<I", 1) produces four bytes in little-endian order with the 1 in the first byte, 01 00 00 00. ">I" is big-endian and gives 00 00 00 01. Both are unsigned 32-bit ints; only the byte order differs.
What is the suffix of Path("a") / "b.tar.gz"?
.suffix is ".gz", because it splits at the last dot, and .stem is "b.tar". Use .suffixes to get both extensions, ['.tar', '.gz'].
Why add ORDER BY to a query whose output is printed?
Without ORDER BY, SQL guarantees no row order, so the same query can return rows in a different order after an insert or a change of index. Printed or judged output must be deterministic, so order explicitly and add tie-breakers.
Exercises
CSV to JSON and back
Read a CSV document (name,qty,price) from standard input with DictReader, convert each row to a dict with qty as int and price as a string, and write the list as JSON (indent=2, sort_keys=True) to items_demo.json using pathlib.Path.write_text. Print the file's size. Read it back with json.loads(path.read_text()), print <name> x<qty> = <qty * price> per item with two decimals (use Decimal for the price), then total <sum>. Delete the file at the end.
Input: a CSV document. Output: size <bytes>, the item lines, total <sum>.
name,qty,price
bolt,3,0.50
nut,10,0.25
prints
size 128
bolt x3 = 1.50
nut x10 = 2.50
total 4.00Binary record store
Write records id score packed as >Ii (big-endian uint32 id, int32 score) to store_demo.bin; print the file size and its SHA-256 hex digest. Then answer get <index> queries by seeking directly to index * record_size and unpacking one record (missing when the index is out of range). Delete the file at the end.
Input: n, then n record lines, then query lines. Output: size <bytes>, sha256 <hex>, then one line per query: <id> <score> or missing.
2
7 -3
9 42
get 1
get 5
prints
size 16
sha256 78f35ff4e3885e4f72f4593cda25699a3c3250086cfedc6ef28e4d03a05c8913
9 42
missingGrouped report with a rollback
Build an in-memory sqlite3 table expenses(person TEXT, category TEXT, cents INTEGER) from add person category cents lines, committing each one immediately with con.commit(). A line batch-fail starts a transaction with with con: in which the next lines up to end are inserted and then a RuntimeError is raised, so the whole batch is rolled back (print rolled back). Finally print <person> <category> <total> rows from a GROUP BY person, category query ordered by person then category, and grand <total>.
Input: lines. Output: rolled back lines as they occur, then the report.
add ada food 500
add ada travel 1200
batch-fail
add bob food 999
end
add bob food 300
prints
rolled back
ada food 500
ada travel 1200
bob food 300
grand 2000In this module: Files and data formats
- Reading and writing files — open, modes, encoding and with
- pathlib — paths as objects
- CSV — reader, writer, DictReader and the quoting rules
- JSON in depth — custom encoders, decoders, dataclasses and config files
- Bytes and binary data — struct, int.to_bytes, base64 and hashlib
- sqlite3 — a SQL database in the standard library
- Checkpoint — Files and data formats (this lesson)
← sqlite3 — a SQL database in the standard library · Type hints in depth — generics, unions, Callable, TypeVar and Protocol →