pathlib: handling file paths in Python

pathlib represents a file path as an object rather than a string, gathering what os.path used to scatter across modules.
5 min read
Believemy logo

Definition

Writing a file path by hand quickly runs into a problem: you have to guess where the folder ends, where the name begins, and which separator the system expects. pathlib answers that annoyance by representing a file path as an object rather than a string. It is a standard library module, whose central class is called Path.

That class gathers what used to be spread across several modules: os.path for the path text, os for the file system, glob for pattern-based search, shutil for copying or moving. A single object now replaces four imports.

Here is what that looks like in practice.

PYTHON
from pathlib import Path

folder = Path("data") / "exports"
report = folder / "sales.csv"

print(report)           # data/exports/sales.csv
print(report.suffix)    # .csv
print(report.stem)      # sales
print(report.exists())  # False

Nothing to install: a single import line is enough, the module has shipped with Python since version 3.4. Born from PEP 428, it was designed as one whole rather than assembled piece by piece across versions.


The problem it solves

A string holding a path is still just a string, nothing more. It does not know where the folder ends, where the extension begins, or whether the file exists. That has to be worked out every single time, with functions scattered across several modules.

Here is what that looks like with the older approach.

PYTHON
import os.path

path = os.path.join(os.path.dirname(__file__), "data", "report.txt")
name = os.path.splitext(os.path.basename(path))[0]

The same operation with a Path object fits on two lines that read almost out loud, with no nesting to untangle.

PYTHON
path = Path(__file__).parent / "data" / "report.txt"
name = path.stem

The gain is more than a matter of fewer characters. Assembling paths with + produces doubled separators or names glued together, and a path hard-coded with forward slashes silently breaks on Windows, where the expected separator is the backslash. That is what the pathlib division operator solves: Path hijacks division through a magic method, the clearest example of that trick in the standard library.


Against os.path

Replacing os.path everywhere would not do the older module justice, though. It is not deprecated, works correctly on strings, and is even a touch faster over millions of repetitions. What it lacks is coherence.

The table below places both ways of writing side by side, for the most common operations.

What you want to doos.path and friendspathlib
Join two piecesos.path.join(a, b)a / b
Step up to the parentos.path.dirname(c)c.parent
The name without extensionos.path.splitext(...)[0]c.stem
Read a whole fileopen(c).read()c.read_text()
Create the folder treeos.makedirs(c)c.mkdir(parents=True)
List the csv filesglob.glob(pattern)c.glob("*.csv")

The fourth row deserves a word of its own. read_text opens, reads and closes the file in one gesture, which saves the with block for a small configuration file. On a large file walked line by line, though, the object method open becomes needed again, and so does the context manager, otherwise the file stays open too long.


The trap for newcomers

The first trap comes from the nature of the object. A Path is not a string, even though it displays like one on screen. Almost the entire standard library accepts one now, and so does pandas. But an older library or a serialisation to json will raise a blunt TypeError. The explicit conversion str(path) settles it.

The second trap is silent, so more costly: if the right-hand side of an assembly is an absolute path, it wipes out everything before it, with no warning.

PYTHON
Path("/var/app") / "logs/app.log"   # /var/app/logs/app.log
Path("/var/app") / "/logs/app.log"  # /logs/app.log
Warning

One extra slash in a configuration value is enough to trigger this trap. The program then writes at the root of the disk, without warning, and the mistake is often discovered much later.

The third trap comes from the reference folder: a relative path resolves against the current directory of the process, never against the location of the code file. The script works launched from its own folder, but fails from a cron job. Path(__file__).resolve().parent anchors the path on the file and puts an end to those bugs.


When it is of no use

pathlib talks about the local file system, and about nothing else. A web address is not a file path, even though it resembles one: handing it to Path squashes the double slash of https, where requests does the job correctly. The same problem touches a key in object storage or an entry inside a zip archive: the visual resemblance is misleading, the logic is not the same.

There is also a question of volume. When walking several hundred thousand entries, every Path object created costs a little memory, and os.scandir stays leaner. For a single path, never assembled, a plain string does the job very well. The most frequent mistake is converting on principle, even where the object brings nothing.


Frequently asked questions

Question

Does all the existing os.path code have to be rewritten?

No, a sweeping rewrite would bring nothing but a risk of regression. Both approaches live side by side without trouble, since the functions of os.path accept a Path object as input. The right moment to switch is when you are already editing that file for another reason.

Question

How can a folder and all its subfolders be walked?

With rglob, the recursive version of glob, inside an ordinary loop. One detail matters: the method hands back a generator rather than a list, so testing it directly to find out whether there are results always answers yes, even when empty.

Question

Why is checking that a file exists before opening it discouraged?

Because the file can disappear between the check and the opening, and the check says nothing about read permissions. Opening directly and catching the exception gives the true answer, the one from the operating system, rather than an answer valid one millisecond earlier.

Related terms

Discover our python glossary

Browse the terms and definitions most commonly used in development with Python.

Share this article

Want to help us? Share this article on your networks or even better: on your site, in an article or in your newsletter.