How to Read and Write Files in Python: open, Modes, with

Python's open takes a mode and an encoding: r reads, w truncates, a appends, b is binary. Reading line by line, writing, and why with matters.

  • Course: Python study plan
  • Module: Files and data formats
  • Kind: Lesson
  • Reading time: 13 min
  • Runtime: CPython 3.11

How do you read a file line by line in Python?

Open the file with with open(path, encoding="utf-8") as f: and loop for line in f:. Iterating the file object reads one line at a time, so even a multi-gigabyte log costs only the memory of its longest line. Each line keeps its newline character; f.read() suits small files read whole, and with closes the file on every exit.

Lesson

open returns a file object, and everything about files in Python follows from three choices made in that call: the mode (read, write, append; text or binary), the encoding (which should always be stated for text), and whether the object is used inside a with block (it should be). This lesson covers those choices, the read and write methods and when each is right, line iteration and the newline rules, print(file=), in-memory files with io.StringIO, and the facts that matter for the judge — a program may create and read files in its working directory — and for real systems, where a file that is not closed is a resource leak and a partly written file is data loss.

open

with open("notes.txt", "w", encoding="utf-8") as f:
    f.write("first line\n")
    f.write("second line\n")

with open("notes.txt", encoding="utf-8") as f:      # mode "r" is the default
    text = f.read()
ModeMeaning
"r"read text (default); FileNotFoundError if absent
"w"write text; truncates an existing file or creates it
"a"append text; creates if absent
"x"create for writing; FileExistsError if present
"r+"read and write, no truncation
"rb", "wb", "ab"the same in binary: bytes in, bytes out

Text mode decodes and encodes with encoding and translates line endings (\r\n on Windows becomes \n on read, and \n becomes the platform's on write — newline="" disables that, which csv requires). Binary mode does neither. State encoding="utf-8" on every text open: the default is the platform's preferred encoding, which is not UTF-8 on every Windows machine, and a file written on one system is then unreadable on another.

with

open acquires a file descriptor — a limited OS resource — and buffers writes in memory. with guarantees close() on every exit (Module 10), which flushes the buffer to disk and releases the descriptor. Without it, a write may sit in the buffer until the interpreter exits, and a loop that opens files without closing them runs out of descriptors. The rule has no exceptions worth learning.

Reading

with open(path, encoding="utf-8") as f:
    everything = f.read()                 # the whole file as one string
    # — or —
    lines = f.readlines()                 # a list of lines, each with its "\n"
    # — or —
    for line in f:                        # one line at a time, lazily — the idiom
        line = line.rstrip("\n")
    # — or —
    first = f.readline()                  # one line; "" at end of file
    chunk = f.read(4096)                  # up to n characters

Iterating the file object is the right default: it reads as it goes and never holds more than one line, so a multi-gigabyte log costs the memory of its longest line. read() is right for small files that are processed as a whole (a config, a template); read(n) in a loop for binary streams. After a full read the position is at the end; f.seek(0) rewinds.

Writing

with open(path, "w", encoding="utf-8") as f:
    f.write("one line\n")                          # write adds no newline
    f.writelines(f"{x}\n" for x in xs)             # an iterable of strings, no separators added
    print("via print", file=f)                     # print adds the newline and str()s its arguments
    print(*values, sep=",", file=f)

write takes a string and returns the count written; it does not add a newline. print(..., file=f) is often the most convenient for line-oriented output because it formats and terminates. Writing is buffered; f.flush() pushes the buffer without closing (for progress files another process reads).

The safe way to replace a file is to write a new one and rename it over the old, so a crash mid-write leaves the original intact: tempfile.NamedTemporaryFile(dir=..., delete=False), write, close, os.replace(tmp, path). Overwriting in place with "w" truncates first and loses the data if the write fails.

Paths, briefly

The path passed to open is relative to the current working directory — wherever the process was started, which is not necessarily the script's directory. Build paths from a known base (pathlib.Path(__file__).parent, next lesson) rather than assuming. Forward slashes work on every platform. On the judge the working directory is writable, and a program may create, write, read and delete files there — which is how the file exercises in this module work; they leave nothing behind.

In-memory files

import io

buf = io.StringIO()
print("captured", file=buf)
buf.getvalue()                    # 'captured\n'

reader = io.StringIO("a,b\n1,2\n")
for line in reader:               # behaves like an open text file
    ...
bio = io.BytesIO(b"\x00\x01")     # the bytes equivalent

StringIO and BytesIO are file objects backed by memory: the way to test code that takes a file, to capture output, and to feed csv/json from a string. Anything that accepts "a file-like object" accepts them.

Errors

FileNotFoundError, PermissionError, IsADirectoryError are subclasses of OSError and carry .filename; UnicodeDecodeError means the encoding was wrong (errors="replace" substitutes U+FFFD, errors="ignore" drops, both hide the problem). EAFP applies: try: open(...) rather than if os.path.exists(...), because the file can vanish between the check and the open.

Pitfalls

  • No encoding= on a text open.
  • A file opened without with — unflushed writes, leaked descriptors, ResourceWarning.
  • "w" when "a" was meant: the file is truncated.
  • read() on a huge file when iteration would do.
  • readlines() then indexing, when a loop was meant; and forgetting the \n on each element.
  • A relative path that depends on the working directory.

Key takeaways

  • open(path, mode, encoding="utf-8") inside with; modes r/w/a/x and b for binary; w truncates.
  • Iterate the file object for lines; read() for small whole files; readline() returns "" at the end.
  • write adds nothing; print(file=f) formats and terminates; writes are buffered until flush/close.
  • Replace files by writing a temporary and os.replace-ing it; paths are relative to the working directory.
  • io.StringIO/BytesIO are in-memory files for tests and captured output; file errors are OSError subclasses — use EAFP.

Common questions

What is the difference between r, w and a modes in Python?

"r" reads text and is the default, raising FileNotFoundError if the file is absent. "w" writes, truncating an existing file to empty first or creating it. "a" appends to the end, creating the file if needed. "x" creates only and fails if the file exists, and adding b makes any mode binary.

Why use with open in Python?

with guarantees the file is closed on every exit, including an exception, and closing flushes buffered writes to disk and releases the operating system's file descriptor. Without it, writes can sit in the buffer and a loop that opens files can run out of descriptors.

Why should I specify the encoding when opening a file in Python?

Without encoding=, text mode uses the platform's preferred encoding, which is not UTF-8 on every Windows machine, so a file written on one system can be unreadable on another. Pass encoding="utf-8" on every text open; a UnicodeDecodeError usually means the encoding was wrong.

How do I safely overwrite a file in Python?

Write the new content to a temporary file in the same directory, close it, then call os.replace(tmp, path) to swap it over the original. If the program crashes mid-write the original is intact, whereas opening the target with "w" truncates it first and loses the data if the write fails.

Exercises

Write, then read back

Read lines from standard input and write them to a file out_demo.txt (UTF-8, one line each, inside a with). Print the file's size in bytes (os.path.getsize). Then reopen it for reading, iterate its lines and print each numbered from 1 with the newline stripped, followed by lines <n>. Delete the file at the end (in a finally).

Input: zero or more lines. Output: bytes <n>, then <i>: <line> per line, then lines <n>.

héllo
world

prints

bytes 13
1: héllo
2: world
lines 2

An append-only log

Maintain app_demo.log from commands: log <text> appends one line using print(..., file=f) on a file opened in append mode; show prints the file's lines numbered; count prints the number of lines by iterating the file; clear truncates it by opening in "w" mode. Delete the file at the end.

Input: commands. Output: the lines of show and count.

log first
log second, with comma
show
count
clear
count

prints

1: first
2: second, with comma
2
0

In this module: Files and data formats

← Checkpoint — The standard library in depth · pathlib — paths as objects →