Definition
Writing age: int in a Python class protects nothing at runtime: nothing stops a caller from passing the string "twenty" instead of an integer. The problem turns real the moment data arrives from outside the program, an API response, a form, a file, because that is exactly where it is most likely to be wrong. Pydantic closes this gap: this third-party library turns a type hint into a check that actually runs, at the moment the data enters the program.
You describe the expected shape as a class, Pydantic takes in a raw dictionary, and it hands back one of two results: an object whose every field carries the right type, or an error naming the offending field and the reason for the refusal. The example below builds a customer from an identifier sent as text.
from pydantic import BaseModel
class Customer(BaseModel):
identifier: int
email: str
active: bool = True
# "42" arrives as text, the way a form would send it
customer = Customer(identifier="42", email="mary@example.com")
print(customer.identifier) # 42, now an integerInstalling it goes through pip, like any library outside the standard library.
pip install pydanticThe problem it solves
The real trouble is not that Python allows this kind of mistake, it is that no tool catches it at the right moment. A static checker such as mypy does catch a contradiction in the code you write, but it never sees what arrives from outside: an API response, a json file, a form filled in by a visitor, a database row.
Here is what a hand-written check looks like, for the first two fields of an object barely richer than the earlier example.
if "identifier" not in data:
raise ValueError("identifier missing")
if not isinstance(data["identifier"], int):
raise TypeError("identifier must be an integer")
if "email" not in data:
raise ValueError("email missing")That block works, but it doubles the size of the function for two fields alone, and a field that changes breaks it. Pydantic replaces all of it with the model declaration: the description of the expected shape becomes the check, and the ValueError or TypeError written one by one stop being necessary.
Pydantic or dataclass
One question comes up quickly: why not just use a dataclass, already shipped with Python? It writes the constructor, the display and the equality of a data object for you, but it never looks at the values it receives. A dataclass accepts "twenty" in a field meant to hold an integer just as readily as it accepts 20.
The table below lists the differences that actually matter for choosing between the two.
| Need | dataclass | Pydantic |
|---|---|---|
| Avoid repetitive code | Yes | Yes |
| Check types at runtime | No | Yes |
| Error message per field | No | Yes |
| Read a nested structure | By hand | Automatic |
| Shipped with Python | Yes | No |
A dataclass suits an object your own code builds; a Pydantic model takes over for an object built from data you did not write.
The three early surprises
Three behaviours almost always surprise someone discovering Pydantic, and it helps to know them before running into them rather than after.
By default, the string "42" placed in an int field triggers no refusal: it becomes the integer 42. That behaviour helps on a form, where everything arrives as text, but it surprises anyone expecting a strict check. Strict mode exists, but it must be asked for explicitly, field by field or model by model.
The second surprise concerns when the check happens: at construction, and nowhere else. Changing an attribute afterwards goes through no check at all, unless validation on assignment has been turned on separately, so a Pydantic object is not watched permanently: it is verified once.
The third has to do with the installed version. Moving from 1 to 2 renamed most of the public interface, and plenty of examples online still speak the old language. The table below matches the two vocabularies.
| Pydantic 1 | Pydantic 2 |
|---|---|
.dict() | .model_dump() |
.json() | .model_dump_json() |
parse_obj() | model_validate() |
@validator | @field_validator |
When it is of no use
On data your program has just produced itself, validating brings nothing: checking twice what you wrote yourself costs time without ever catching an error. In a loop building hundreds of thousands of objects, that cost becomes measurable, and a plain dataclass takes the lead again.
It does not replace a static checker either, since that one works before launch on code Pydantic never sees, and it replaces tests even less: a respected model can describe a negative price if nobody asked for it to be positive. Validating a shape is not validating a business rule, unless that rule is written down explicitly in the model.
Frequently asked questions
Does Pydantic replace mypy?
No, the two work at different moments. mypy reads the code before execution; Pydantic reads values during execution and refuses the ones that do not match. Serious projects use both, on the very same hints.
How can the errors be caught instead of letting the program crash?
By wrapping the model construction in a try block and catching ValidationError with an except. That exception exposes an errors() method returning the list of refused fields, each with its path and its reason: enough to build a readable response rather than let a crash trace bubble up.
Does every dictionary in a project need a model?
No, and that is the most common excess among people discovering the library. A model belongs at the entrances and exits of the program, where data changes hands; in between, the code works on objects already validated. Hence its place inside FastAPI, where the model acts as both check and interface documentation.