Python pathlib vs os.path: Working with Paths as Objects
Python's pathlib models a path as an object: join with a slash, split name, stem and suffix, list with glob and rglob, and read, write and create files.
- Course: Python study plan
- Module: Files and data formats
- Kind: Lesson
- Reading time: 13 min
- Runtime: CPython 3.11
What is pathlib in Python?
pathlib is the standard-library module that represents file-system paths as Path objects instead of strings. Path("data") / "q1.csv" joins parts on every platform; .name, .stem, .suffix and .parent take a path apart; .exists(), .read_text() and .glob() work with the disk. It replaces the string functions of os.path in new code.
Lesson
A path is not a string. It has a directory part, a name, a stem and a suffix; it can be joined, made absolute, tested for existence, listed, created and removed; and the separator differs between operating systems. pathlib.Path (3.4) models all of that as an object with methods, replacing the string juggling of os.path with / for joining and readable names for everything else. This lesson covers constructing and joining paths, the parts, the queries, reading and writing through a path, listing and searching a tree, creating and deleting, and the pure-path classes that manipulate paths without touching the file system.
Constructing and joining
from pathlib import Path
p = Path("data") / "reports" / "q1.csv" # / joins, on every platform
p # PosixPath('data/reports/q1.csv') (WindowsPath on Windows)
str(p) # 'data/reports/q1.csv' — for APIs that want a string
Path.cwd(), Path.home() # the working directory, the home directory
Path(__file__).parent # the directory this script lives in
p.resolve() # absolute, symlinks resolved
Path("a/b/../c") # not simplified; resolve() does that
Path("dir") / "file" is the idiom; the operator takes strings or paths on the right. Paths are immutable and hashable, so they work as dict keys and set members. Nearly every function that takes a filename (open, os.remove, shutil.copy, json.load via open) accepts a Path.
The parts
p = Path("/home/ada/data/report.tar.gz")
p.name # 'report.tar.gz'
p.stem # 'report.tar'
p.suffix # '.gz'
p.suffixes # ['.tar', '.gz']
p.parent # PosixPath('/home/ada/data')
p.parents[1] # PosixPath('/home/ada')
p.parts # ('/', 'home', 'ada', 'data', 'report.tar.gz')
p.anchor # '/'
p.with_suffix(".zip") # PosixPath('/home/ada/data/report.tar.zip') — replaces the LAST suffix
p.with_name("summary.txt")
p.with_stem("report2") # 3.9
p.is_absolute()
p.relative_to("/home/ada") # PosixPath('data/report.tar.gz'); ValueError if not under it
stem and suffix split at the last dot, which is right for report.csv and surprising for report.tar.gz; suffixes has them all.
Queries
p.exists(), p.is_file(), p.is_dir(), p.is_symlink()
p.stat().st_size # bytes
p.stat().st_mtime # modification time (a timestamp)
p.samefile(other)
These touch the file system; the ones above did not. As always, a check followed by an action has a window; prefer trying the action and catching OSError when the outcome matters.
Reading and writing through a path
text = p.read_text(encoding="utf-8") # the whole file
p.write_text("hello\n", encoding="utf-8") # creates or truncates
data = p.read_bytes(); p.write_bytes(b"...")
with p.open(encoding="utf-8") as f: # the same as open(p, ...)
for line in f:
...
read_text/write_text are the two-liners for small files; open remains the tool for streaming. State the encoding here too.
Listing and searching
for child in sorted(Path("data").iterdir()): # direct children, files and directories — sort for a defined order
print(child.name)
list(Path("data").glob("*.csv")) # matching names in one directory
list(Path("data").rglob("*.csv")) # recursively, every depth
list(Path("data").glob("**/*.csv")) # the same as rglob
[p for p in Path("src").rglob("*.py") if "test" not in p.parts]
iterdir, glob and rglob yield in file-system order, which is arbitrary — sort before printing or relying on it. Patterns are shell-style (*, ?, [abc]), and rglob walks directories, which can be slow on large trees; os.walk is the lower-level alternative that yields (dirpath, dirnames, filenames) and lets you prune dirnames in place.
Creating, moving, deleting
Path("out/logs").mkdir(parents=True, exist_ok=True) # like mkdir -p
p.touch() # create empty or update mtime
p.rename("new_name.txt") # or replace() to overwrite an existing target
p.unlink(missing_ok=True) # delete a file (3.8 for missing_ok)
Path("empty_dir").rmdir() # only if empty
import shutil
shutil.rmtree("tree") # delete recursively — no undo
shutil.copy(src, dst); shutil.move(src, dst)
mkdir(parents=True, exist_ok=True) is the form that never raises for an existing directory; rmtree is the one call to think twice about.
Pure paths
PurePosixPath and PureWindowsPath do the parsing, joining and parts without ever touching a disk — for manipulating paths for another system, or in tests that must not depend on the file system:
from pathlib import PurePosixPath, PureWindowsPath
PurePosixPath("/srv/app") / "logs" / "x.log" # PurePosixPath('/srv/app/logs/x.log')
PureWindowsPath(r"C:\Users\ada").parts # ('C:\\', 'Users', 'ada')
os.path, for reading old code
os.path.join(a, b), os.path.basename, dirname, splitext, exists, abspath are the string-based ancestors; each has a Path equivalent (/, .name, .parent, .stem/.suffix, .exists(), .resolve()). New code uses pathlib; os.path still appears everywhere and reads easily once the mapping is known.
A worked example: renaming by pattern
for path in sorted(Path("photos").glob("IMG_*.jpg")):
number = path.stem.removeprefix("IMG_")
target = path.with_name(f"holiday_{int(number):04d}{path.suffix}")
if not target.exists():
path.rename(target)
Every step is a path operation: glob finds the candidates, stem and suffix take the name apart, with_name builds the new one beside the old, exists guards against clobbering, rename moves. No string slicing, no separators, and the same code runs on every platform.
Pitfalls
- Building paths with
+and hard-coded slashes. p.suffixon a double extension.- Relying on the order of
iterdir/glob— sort. mkdirwithoutparents=True, exist_ok=Truein a setup step that may run twice.rmtreeon a path built from user input without checking what it resolves to.- Passing a
Pathto a library that insists onstr—str(p)oros.fspath(p).
Key takeaways
Path(...) / "part"joins;.name,.stem,.suffix,.parent,.parts,.with_suffixdecompose and rebuild..exists(),.is_file(),.stat()query;.read_text/.write_text(with an encoding) and.open()do I/O.iterdir,glob,rgloblist — sort the results;mkdir(parents=True, exist_ok=True),unlink,rename,shutil.rmtreechange the tree.- Pure paths manipulate without touching the disk.
- Paths are immutable, hashable and accepted wherever a filename is.
Common questions
How do I join paths in Python?
With pathlib, use the / operator: Path("data") / "reports" / "q1.csv" inserts the right separator on every platform and accepts strings or paths on the right. The older equivalent is os.path.join("data", "reports", "q1.csv"). Never build paths by concatenating strings with hard-coded slashes.
How do I get a file extension in Python?
Path(name).suffix returns the last extension, such as .gz, and .stem the name without it. Both split at the last dot, so for report.tar.gz the suffix is .gz and the stem report.tar; .suffixes returns every extension, ['.tar', '.gz'], and with_suffix(".zip") replaces the last one.
How do I list all files in a directory recursively in Python?
Path("data").rglob("*.csv") yields every matching path at every depth, glob("*.csv") searches one directory and iterdir() lists direct children. All three yield in arbitrary file-system order, so wrap them in sorted() when order matters. os.walk is the lower-level alternative that lets you prune directories.
How do I create a directory if it does not exist in Python?
Call Path("out/logs").mkdir(parents=True, exist_ok=True), which creates any missing parents, like mkdir -p, and does nothing if the directory already exists. Without those arguments it raises FileNotFoundError for a missing parent and FileExistsError for an existing directory.
What is the difference between pathlib and os.path?
os.path works on strings with functions such as join, basename, dirname and splitext; pathlib offers the same operations on an immutable Path object (/, .name, .parent, .stem, .suffix) and adds I/O such as read_text. New code uses pathlib, and most functions that take a filename accept either.
Exercises
Path parts
Read POSIX path strings, one per line, and decompose each with PurePosixPath (no file system involved): print name, stem, suffix, all suffixes joined with + (or - if none), parent, and the path with its suffix replaced by .bak.
Input: lines. Output: name=<n> stem=<s> suffix=<x> suffixes=<a+b> parent=<p> bak=<q> per line.
/home/ada/data/report.tar.gz
prints
name=report.tar.gz stem=report.tar suffix=.gz suffixes=.tar+.gz parent=/home/ada/data bak=/home/ada/data/report.tar.bakBuild a tree and search it
Create a directory tree_demo next to this script and populate it from mk <relative/path> lines (creating parent directories with mkdir(parents=True, exist_ok=True) and empty files with touch). Then for each find <pattern> line print the matching paths from rglob relative to the root, sorted, one per line (or none), and for count print the number of files in the tree. Remove the tree at the end with shutil.rmtree.
Input: commands. Output: the results of find and count.
mk docs/a.md
mk docs/sub/b.md
mk src/main.py
find *.md
count
prints
docs/a.md
docs/sub/b.md
3In this module: Files and data formats
- Reading and writing files — open, modes, encoding and with
- pathlib — paths as objects (this lesson)
- CSV — reader, writer, DictReader and the quoting rules
- JSON in depth — custom encoders, decoders, dataclasses and config files
- Bytes and binary data — struct, int.to_bytes, base64 and hashlib
- sqlite3 — a SQL database in the standard library
- Checkpoint — Files and data formats
← Reading and writing files — open, modes, encoding and with · CSV — reader, writer, DictReader and the quoting rules →