Definition
A type hint is a promise. Nothing in the language checks that it holds: writing total: float commits to nothing more than a comment, as long as nobody rereads it against what the code actually does with the value. That gap is what mypy fills. It reads the code without ever running it, compares what a type hint announces against how the value is used further down, and reports every place where the two disagree, without rewriting anything or altering the program's behaviour. It installs with pip and runs from the command line.
python -m mypy my_project/Without it, the inconsistency simply waits its turn: it only surfaces when an operation fails on its own account, often much further down the program, in the shape of a TypeError. mypy brings that moment forward by several hours, sometimes by several weeks, and pins it back to the exact line where the mistake was written.
# mypy compares the "float" hint against the real call right below it
def unit_price(total: float, quantity: int) -> float:
return total / quantity
unit_price("120", 4)
# error: Argument 1 has incompatible type "str"; expected "float"The problem it actually solves
It is not the mistake in the example above that a checker spends most of its time catching on real code: a string handed over where a number was expected shows up on the first run, with no tool needed to flag it. What slips past the eye instead is the missing value that travels across several functions before failing. A function able to return None, ten places calling it, and a single one forgetting to check before using the result.
# dict | None warns that the result may be missing; the call below ignores that
def find_customer(identifier: int) -> dict | None:
return database.read(identifier)
customer = find_customer(12)
print(customer["name"])
# error: Value of type "dict | None" is not indexableThat is the bulk of what a checker reports on code already running in production: not exotic types, but unhandled None values, an attribute renamed six months ago, a parameter dropped from a function that a distant caller still passes. Every one of them a bug that would otherwise have announced itself later as a TypeError or an AttributeError, at the worst possible hour.
What it cannot see
A checker only knows what is written down in black and white. On a project with no hints at all, mypy prints "Success: no issues found", and that sentence says nothing reassuring: it checked nothing, having no promise to compare against.
Reading that message as a green light is the costliest mistake anyone can make with mypy: on a project with no hints, it proves nothing at all, and the bug you thought was ruled out is simply waiting its turn.
The tool is just as blind to what arrives from outside the program. A JSON file, an API response, or a database row all turn up with no known type, and the hint written over them is taken at face value rather than checked. Confirming the real shape of data on arrival falls to a validation library such as Pydantic, not to mypy. Finally, it reasons about declarations rather than values: a function returning a wrong price with the right type goes through without a word.
Each defect has its own tool, and mixing them up sends you looking in the wrong place:
| The defect | The tool that sees it |
|---|---|
| An unused import, a variable never read again | ruff |
| Quotes and line breaks that change from one file to the next | black |
A None passed where a string is expected | mypy |
| A calculation right on paper and wrong on real data | pytest |
| A field missing from the JSON received this morning | Pydantic |
The classic beginner mistake
It consists of running strict mode over a project that has existed for years. The terminal prints nine hundred errors at once, nearly all of them saying only that some function carries no hints, and the tool gets uninstalled within the half hour. Not one of those messages points at a bug: they describe the absence of hints, which was already known before the command was typed.
The approach that holds up over time does the opposite. One file at a time, the one being edited anyway for some other reason, annotated along the way, and strict mode kept for new modules, switched on folder by folder in the project configuration.
python -m mypy billing/invoices.py
python -m mypy --strict billing/new_module.pyA second trap lies in wait, the same one as for any Python package: installed outside the project virtual environment, the checker cannot see the libraries living there and asks for types it will never find. The python -m mypy form, rather than bare mypy, rules out that risk. As for the message announcing missing types for some library, it calls for installing a separate package, usually named types-something, not for slapping on an ignore comment.
mypy, Pyright, or nothing at all
The direct alternative is called Pyright, and plenty already use it without knowing: it is the engine underlining their code inside Visual Studio Code. Both tools do the same underlying job, with different trade-offs. The table below sums up what separates them:
| Criterion | What separates them |
|---|---|
| Speed | Pyright stays clearly faster on a large project |
| Default severity | Pyright infers and reports more, mypy expects explicit hints |
| When feedback lands | Pyright underlines as you type, mypy runs on demand |
| Ecosystem | mypy is the reference implementation of the typing PEP documents |
The real advice fits on one line: pick a single one for the whole team, or it ends up arguing about contradictory messages on the edge cases instead of fixing code. There are also projects where neither is worth the trouble: a forty-line script thrown away after use, a throwaway analysis notebook. The benefit comes from how long the code lives and how many hands touch it, not from the size of the file. On very dynamic code leaning openly on duck typing, writing hints even becomes tedious for what it gives back.
Frequently asked questions
Does the whole project need hints before mypy is worth anything?
No, and waiting until everything is annotated is the surest way never to start. A project annotated only on its most called functions already pays off, since type errors travel along those very functions. One reservation is worth knowing though: by default it skips the body of unannotated functions, so as long as a signature stays bare, nothing will be said about what happens inside it.
Does it replace tests?
No, the two are not looking for the same thing. A checker proves that the pieces fit together, a test verifies that the answer is the right one. A billing function returning a wrong amount with the right type goes through without a remark and fails on the first serious test: neither one covers the other's blind spot.
What should be done when the checker is wrong?
It happens, mostly around poorly typed libraries. The clean way out is placing an ignore comment on the single line concerned, carrying the precise error code in brackets and a sentence explaining why. What costs dearly is the ignore applied to a whole file: it switches checking off for good, and nobody notices until the bug arrives.