Definition
Something constantly has to be found inside a text without its value being known in advance: the date in the middle of a log line, the number in an email subject, the postcode lost in an address. Looking for a known word raises no difficulty, the in operator is enough. Here, though, what is being looked for is unknown: only its shape is known.
That is the gap a regular expression fills. It is a pattern: a short formula describing the shape of a text rather than its content. "Four digits, a dash, two digits, a dash, two digits" is a pattern, and 2026-08-19 is one of the texts that match it.
Python keeps these patterns in the module called re, shipped with the standard library: nothing to install with pip, a single import re at the top of the file is enough.
import re
text = "Order 4821 shipped on 2026-08-19"
# Four digits, a dash, two digits, a dash, two digits
result = re.search(r"\d{4}-\d{2}-\d{2}", text)
print(result.group()) # 2026-08-19Everything happens inside the string passed as the first argument. \d means "a digit", {4} means "four times in a row", and the dash means nothing special: it is looked for as it is. A pattern therefore constantly mixes characters taken literally with symbols describing a family or a repetition.
That syntax does not belong to Python. Barring a few variations, the same one is written in JavaScript, in PHP, in grep or in the search field of a code editor: what is learned here will serve elsewhere for years.
What it replaces, and what it must not replace
A regular expression replaces one precise thing: the tests and the manual splits written to catch an element whose position changes from one line to the next. Pulling an invoice number out of an email subject, collecting the postcodes in an export, checking that a product reference has the expected shape.
Without a pattern, that work is done with nested split calls, and the code turns unreadable at the second exception met in the data: a line where the space is missing, another where the field is empty. The pattern absorbs those variations because it describes a shape and not a position.
It does not, on the other hand, replace string methods, which on a fixed need read better and run faster. Here is how the most common cases divide up.
| The need | The right tool |
|---|---|
| Knowing whether a word appears in a text | The in operator |
| Splitting on a fixed separator | split |
| Checking a beginning or an end | startswith, endswith |
| Finding a variable shape | A regular expression |
| Reading HTML, JSON or CSV | The format parser, never a pattern |
That last row is the one that costs the most when ignored. HTML, JSON and CSV allow nesting, escaped quotes and line breaks in the middle of a value: a pattern that passes on three examples breaks on the fourth real file, once the data has already been processed.
The json module, the csv module or an HTML parser know the rules of the format and apply all of them, in a single line. And if the data already sits in a pandas table, .str.extract applies a pattern to a whole column: a pattern works on values, never on a structure.
The functions of the re module
The module exposes a dozen functions, but seven cover nearly every use. They all take the pattern as their first argument and the text as their second; what separates them is what they hand back, so the right-hand column.
| Call | What it returns |
|---|---|
re.search | The first match found anywhere, or None |
re.match | The same, but only at the start of the text |
re.fullmatch | A result only if the whole text matches |
re.findall | The list of every match found |
re.finditer | An iterator over the full result objects |
re.sub | The text with matches replaced |
re.split | The text split on the pattern |
The first two rows are the leading cause of lost time for anyone discovering the library. re.match anchors the search at the very first character: looking for "cat" with match inside "the cat sleeps" returns nothing, and a quarter of an hour goes into repairing a pattern that has no fault.
Two reflexes are enough. To find something inside a text, re.search. To check that a whole text has the right shape, a postcode typed into a form for instance, re.fullmatch. re.match only earns its place where the beginning of the text genuinely matters.
These functions all accept a flags argument: re.IGNORECASE ignores case, re.MULTILINE makes ^ and $ apply to each line rather than to the whole text, and re.VERBOSE allows spaces and comments inside the pattern, the only way to keep a long pattern readable.
10 matches. The highlighted parts are what findall would hand back. Groups in brackets are extracted separately, which is how you pull out part of what you found.
The engine here is the browser's, very close to Python's on common patterns but not identical. Python names its groups (?P<name>...) where the browser writes (?<name>...), and the re module offers flags this demonstration does not expose.
The result is not the text found
A trap waits right after the first successful search, and it only goes off in production. A search does not return a string but a match object, carrying the text found, its position and its groups. And when the pattern finds nothing, it returns None.
Chaining .group() onto that return therefore works as long as the text holds what is expected. The day an email arrives without an invoice number, it is None that receives the call: an AttributeError is raised and the script stops. The test that avoids it fits on one line.
result = re.search(r"invoice (\d+)", subject)
if result: # the test that is almost always missing
number = result.group(1) # what the parentheses captured
else:
number = NoneParentheses play a second role here: they mark a group, the piece of the result to be picked up on its own. group(1) returns what the first pair captured, group(0) the whole match. Past two or three groups, rely on names instead: (?P<number>\d+) reads without counting parentheses and feeds groupdict(), which hands back a ready-made dictionary.
One writing detail that is not one: the r prefix in front of the pattern. Without it, Python processes the backslashes on its own account before the engine ever sees them, and the pattern arrives damaged. Writing r"..." every time wipes out a whole family of baffling bugs.
Greedy by default
One last behaviour of the engine explains most of the patterns that "catch too much". The * and + quantifiers are greedy: they swallow as much text as they can, then give back just enough for the end of the pattern to match.
line = 'name="Dupont" city="Lyon"'
re.findall(r'".*"', line) # greedy : ['"Dupont" city="Lyon"']
re.findall(r'".*?"', line) # lazy : ['"Dupont"', '"Lyon"']The first pattern does not stop at the quote closing Dupont: it runs to the last quote on the line. The question mark makes the second one lazy, meaning it stops at the first opportunity, and the two values come out separately.
An often better writing forbids the ending character outright, here r'"[^"]*"': the pattern can no longer overflow and the engine has no backtracking to do. On a long text, nested quantifiers such as (a+)+ take a program from a few milliseconds to several minutes, and if that text comes from a public form, a single request is enough to freeze the server.
Frequently asked questions
Why does my pattern work in an online tester but not in my code?
Three causes, in order of frequency. The r prefix was forgotten, and the pattern reaches the engine damaged. Or re.match was used where re.search was meant, and the search stays glued to the first character. Or finally the options testers switch on by default, global search and multiline mode, are not passed explicitly in Python.
Should patterns be compiled with re.compile?
Rarely for speed: the module caches the patterns it has just used, so the gain stays invisible inside an ordinary loop. re.compile mainly serves to give the pattern a name and to place it once and for all at the top of the file, together with its options, instead of repeating it in three places. It is a readability choice first.
How can an email address be validated with a regular expression?
As honestly as possible, which means barely. The specification allows shapes nobody ever writes, and the pattern covering them all runs to several thousand characters without being correct anyway. In practice, check that there is a single at sign, text on both sides and a dot after it, then send a confirmation message. Only the message that arrives proves the address exists.