Python pathlib vs os.path: Working with Paths as Objects

Python's pathlib models a path as an object: join with a slash, split name, stem and suffix, list with glob and rglob, and read, write and create files.

  • Course: Python study plan
  • Module: Files and data formats
  • Kind: Lesson
  • Reading time: 13 min
  • Runtime: CPython 3.11

What is pathlib in Python?

pathlib is the standard-library module that represents file-system paths as Path objects instead of strings. Path("data") / "q1.csv" joins parts on every platform; .name, .stem, .suffix and .parent take a path apart; .exists(), .read_text() and .glob() work with the disk. It replaces the string functions of os.path in new code.

Lesson

A path is not a string. It has a directory part, a name, a stem and a suffix; it can be joined, made absolute, tested for existence, listed, created and removed; and the separator differs between operating systems. pathlib.Path (3.4) models all of that as an object with methods, replacing the string juggling of os.path with / for joining and readable names for everything else. This lesson covers constructing and joining paths, the parts, the queries, reading and writing through a path, listing and searching a tree, creating and deleting, and the pure-path classes that manipulate paths without touching the file system.

Constructing and joining

from pathlib import Path

p = Path("data") / "reports" / "q1.csv"       # / joins, on every platform
p                                             # PosixPath('data/reports/q1.csv') (WindowsPath on Windows)
str(p)                                        # 'data/reports/q1.csv' — for APIs that want a string
Path.cwd(), Path.home()                       # the working directory, the home directory
Path(__file__).parent                         # the directory this script lives in
p.resolve()                                   # absolute, symlinks resolved
Path("a/b/../c")                              # not simplified; resolve() does that

Path("dir") / "file" is the idiom; the operator takes strings or paths on the right. Paths are immutable and hashable, so they work as dict keys and set members. Nearly every function that takes a filename (open, os.remove, shutil.copy, json.load via open) accepts a Path.

The parts

p = Path("/home/ada/data/report.tar.gz")
p.name          # 'report.tar.gz'
p.stem          # 'report.tar'
p.suffix        # '.gz'
p.suffixes      # ['.tar', '.gz']
p.parent        # PosixPath('/home/ada/data')
p.parents[1]    # PosixPath('/home/ada')
p.parts         # ('/', 'home', 'ada', 'data', 'report.tar.gz')
p.anchor        # '/'
p.with_suffix(".zip")        # PosixPath('/home/ada/data/report.tar.zip') — replaces the LAST suffix
p.with_name("summary.txt")
p.with_stem("report2")       # 3.9
p.is_absolute()
p.relative_to("/home/ada")   # PosixPath('data/report.tar.gz'); ValueError if not under it

stem and suffix split at the last dot, which is right for report.csv and surprising for report.tar.gz; suffixes has them all.

Queries

p.exists(), p.is_file(), p.is_dir(), p.is_symlink()
p.stat().st_size                 # bytes
p.stat().st_mtime                # modification time (a timestamp)
p.samefile(other)

These touch the file system; the ones above did not. As always, a check followed by an action has a window; prefer trying the action and catching OSError when the outcome matters.

Reading and writing through a path

text = p.read_text(encoding="utf-8")          # the whole file
p.write_text("hello\n", encoding="utf-8")     # creates or truncates
data = p.read_bytes(); p.write_bytes(b"...")
with p.open(encoding="utf-8") as f:           # the same as open(p, ...)
    for line in f:
        ...

read_text/write_text are the two-liners for small files; open remains the tool for streaming. State the encoding here too.

Listing and searching

for child in sorted(Path("data").iterdir()):     # direct children, files and directories — sort for a defined order
    print(child.name)
list(Path("data").glob("*.csv"))                  # matching names in one directory
list(Path("data").rglob("*.csv"))                 # recursively, every depth
list(Path("data").glob("**/*.csv"))               # the same as rglob
[p for p in Path("src").rglob("*.py") if "test" not in p.parts]

iterdir, glob and rglob yield in file-system order, which is arbitrary — sort before printing or relying on it. Patterns are shell-style (*, ?, [abc]), and rglob walks directories, which can be slow on large trees; os.walk is the lower-level alternative that yields (dirpath, dirnames, filenames) and lets you prune dirnames in place.

Creating, moving, deleting

Path("out/logs").mkdir(parents=True, exist_ok=True)   # like mkdir -p
p.touch()                                              # create empty or update mtime
p.rename("new_name.txt")                               # or replace() to overwrite an existing target
p.unlink(missing_ok=True)                              # delete a file (3.8 for missing_ok)
Path("empty_dir").rmdir()                              # only if empty
import shutil
shutil.rmtree("tree")                                  # delete recursively — no undo
shutil.copy(src, dst); shutil.move(src, dst)

mkdir(parents=True, exist_ok=True) is the form that never raises for an existing directory; rmtree is the one call to think twice about.

Pure paths

PurePosixPath and PureWindowsPath do the parsing, joining and parts without ever touching a disk — for manipulating paths for another system, or in tests that must not depend on the file system:

from pathlib import PurePosixPath, PureWindowsPath
PurePosixPath("/srv/app") / "logs" / "x.log"          # PurePosixPath('/srv/app/logs/x.log')
PureWindowsPath(r"C:\Users\ada").parts                 # ('C:\\', 'Users', 'ada')

os.path, for reading old code

os.path.join(a, b), os.path.basename, dirname, splitext, exists, abspath are the string-based ancestors; each has a Path equivalent (/, .name, .parent, .stem/.suffix, .exists(), .resolve()). New code uses pathlib; os.path still appears everywhere and reads easily once the mapping is known.

A worked example: renaming by pattern

for path in sorted(Path("photos").glob("IMG_*.jpg")):
    number = path.stem.removeprefix("IMG_")
    target = path.with_name(f"holiday_{int(number):04d}{path.suffix}")
    if not target.exists():
        path.rename(target)

Every step is a path operation: glob finds the candidates, stem and suffix take the name apart, with_name builds the new one beside the old, exists guards against clobbering, rename moves. No string slicing, no separators, and the same code runs on every platform.

Pitfalls

  • Building paths with + and hard-coded slashes.
  • p.suffix on a double extension.
  • Relying on the order of iterdir/glob — sort.
  • mkdir without parents=True, exist_ok=True in a setup step that may run twice.
  • rmtree on a path built from user input without checking what it resolves to.
  • Passing a Path to a library that insists on str — str(p) or os.fspath(p).

Key takeaways

  • Path(...) / "part" joins; .name, .stem, .suffix, .parent, .parts, .with_suffix decompose and rebuild.
  • .exists(), .is_file(), .stat() query; .read_text/.write_text (with an encoding) and .open() do I/O.
  • iterdir, glob, rglob list — sort the results; mkdir(parents=True, exist_ok=True), unlink, rename, shutil.rmtree change the tree.
  • Pure paths manipulate without touching the disk.
  • Paths are immutable, hashable and accepted wherever a filename is.

Common questions

How do I join paths in Python?

With pathlib, use the / operator: Path("data") / "reports" / "q1.csv" inserts the right separator on every platform and accepts strings or paths on the right. The older equivalent is os.path.join("data", "reports", "q1.csv"). Never build paths by concatenating strings with hard-coded slashes.

How do I get a file extension in Python?

Path(name).suffix returns the last extension, such as .gz, and .stem the name without it. Both split at the last dot, so for report.tar.gz the suffix is .gz and the stem report.tar; .suffixes returns every extension, ['.tar', '.gz'], and with_suffix(".zip") replaces the last one.

How do I list all files in a directory recursively in Python?

Path("data").rglob("*.csv") yields every matching path at every depth, glob("*.csv") searches one directory and iterdir() lists direct children. All three yield in arbitrary file-system order, so wrap them in sorted() when order matters. os.walk is the lower-level alternative that lets you prune directories.

How do I create a directory if it does not exist in Python?

Call Path("out/logs").mkdir(parents=True, exist_ok=True), which creates any missing parents, like mkdir -p, and does nothing if the directory already exists. Without those arguments it raises FileNotFoundError for a missing parent and FileExistsError for an existing directory.

What is the difference between pathlib and os.path?

os.path works on strings with functions such as join, basename, dirname and splitext; pathlib offers the same operations on an immutable Path object (/, .name, .parent, .stem, .suffix) and adds I/O such as read_text. New code uses pathlib, and most functions that take a filename accept either.

Exercises

Path parts

Read POSIX path strings, one per line, and decompose each with PurePosixPath (no file system involved): print name, stem, suffix, all suffixes joined with + (or - if none), parent, and the path with its suffix replaced by .bak.

Input: lines. Output: name=<n> stem=<s> suffix=<x> suffixes=<a+b> parent=<p> bak=<q> per line.

/home/ada/data/report.tar.gz

prints

name=report.tar.gz stem=report.tar suffix=.gz suffixes=.tar+.gz parent=/home/ada/data bak=/home/ada/data/report.tar.bak

Build a tree and search it

Create a directory tree_demo next to this script and populate it from mk <relative/path> lines (creating parent directories with mkdir(parents=True, exist_ok=True) and empty files with touch). Then for each find <pattern> line print the matching paths from rglob relative to the root, sorted, one per line (or none), and for count print the number of files in the tree. Remove the tree at the end with shutil.rmtree.

Input: commands. Output: the results of find and count.

mk docs/a.md
mk docs/sub/b.md
mk src/main.py
find *.md
count

prints

docs/a.md
docs/sub/b.md
3

In this module: Files and data formats

← Reading and writing files — open, modes, encoding and with · CSV — reader, writer, DictReader and the quoting rules →