The Fsync That Lied on Retry — Python Bug Hunt

Modelled on PostgreSQL's "fsyncgate" (2018): developers found that when fsync failed, Linux could report the error once, mark the dirty pages clean and drop…

  • Language: Python
  • Layer: Database
  • Difficulty: Hard
  • Concepts: Retries, State
  • Modelled on: PostgreSQL · fsyncgate 2018
  • Visible tests: a clean checkpoint makes pages durable and truncates the WAL; a failed fsync panics and is never retried
  • Reward: 50 XP for a complete fix

Briefing

Modelled on PostgreSQL's "fsyncgate" (2018): developers found that when fsync failed, Linux could report the error once, mark the dirty pages clean and drop them. PostgreSQL's checkpointer retried the fsync, the retry succeeded, and the data it believed durable had never reached disk. The fix, released in February 2019, was to treat a failed fsync as fatal: PANIC, and recover from the write-ahead log on restart.

This project is a reconstruction with a fake OS that behaves the way Linux did. checkpoint.py writes buffers, fsyncs, and truncates the WAL — retrying once when the fsync fails.

Fix checkpoint so a failed fsync is never retried and never trusted.

Bug report

BUG-FSYNCGATE · Priority: Critical (silent data loss) · Reported by: storage team

checkpoint(os_, buffers, wal) — wal is {"records": [...], "redo_lsn": int}:

  • writes every buffer (sorted by page name) with os_.write, then calls os_.fsync() exactly once
  • on success: redo_lsn += number of records, records = [], return that number
  • on ANY failure of fsync: raise errors.Panic — do not call fsync again, and leave the WAL exactly as it was, so crash recovery (recovery.recover) can replay it

Observed: after an EIO the checkpoint retried, succeeded, truncated the WAL, and the pages were nowhere on disk.

Logs

[checkpointer] fsync failed: EIO, retrying
[checkpointer] fsync ok, checkpoint complete redo_lsn=412
[startup] page p1 missing, no WAL to replay

The code as shipped

src/storage/checkpoint.py (editable)

errors = bug_require("./errors.py")


def checkpoint(os_, buffers, wal):
    for page in sorted(buffers):
        os_.write(page, buffers[page])
    try:
        os_.fsync()
    except OSError:
        os_.fsync()
    flushed = len(wal["records"])
    wal["records"] = []
    wal["redo_lsn"] += flushed
    return flushed

Read-only context: src/storage/errors.py, src/storage/fakeos.py, src/storage/recovery.py.

Open the hunt to edit the files, run the visible tests and submit against the hidden ones. More Python bug hunts.