Half an Emoji in the Username — Python Bug Hunt

Inspired by MySQL's legendary utf8-that-isn't (3 bytes max — real UTF-8 needs utf8mb4), and the truncation bugs that slice multi-byte characters in half on…

  • Language: Python
  • Layer: Database
  • Difficulty: Medium
  • Concepts: Unicode, Encoding
  • Modelled on: MySQL utf8
  • Visible tests: ascii within budget is untouched; the result never exceeds the byte budget; multi-byte characters are kept whole
  • Reward: 50 XP for a complete fix

Briefing

Inspired by MySQL's legendary utf8-that-isn't (3 bytes max — real UTF-8 needs utf8mb4), and the truncation bugs that slice multi-byte characters in half on the way into a column.

truncate.py fits a username into a byte budget. It must never split a character.

Bug report

BUG-MB4 · Priority: High · Reported by: storage

fit_bytes(text, max_bytes):

  • the UTF-8 encoding of the result is at most max_bytes long
  • characters are kept whole — a 4-byte emoji that doesn't fit is dropped
  • keep the longest valid prefix

Observed: slicing by characters overflows the byte budget, and the write path then hard-truncates mid-emoji, corrupting the row.

Logs

[db] value for column 'display_name' exceeds byte limit; hard-truncated
[render] username shows U+FFFD replacement character

The code as shipped

src/db/truncate.py (editable)

# Fits text into a UTF-8 byte budget without corrupting characters.

def fit_bytes(text, max_bytes):
    return text[:max_bytes]

Open the hunt to edit the files, run the visible tests and submit against the hidden ones. More Python bug hunts.