Definition
Knowing whether a value sits somewhere is one of the most ordinary questions in the job. Is this username already taken? Is this address on the blocklist? In many languages, answering it takes a loop, a flag and an early exit: five lines for a question that is worth one.
Python made the opposite choice and set a keyword aside for that question. in asks it directly, and it returns nothing but a true or false answer, never a position nor the item found. That is what makes it a condition ready to use as it is behind an if, with no intermediate variable.
languages = ["python", "javascript", "go"]
if "python" in languages:
print("Present")Read aloud, the line is almost a plain English sentence. That readability explains why in almost always replaces the search loop written by hand, along with its off-by-one mistakes and its forgotten break statements.
Two roles for one word
An awkwardness turns up quickly, with the very first loop written. The same keyword comes back in the for loop, where it tests nothing at all: it separates the variable from the sequence being walked through, the way punctuation would. The two uses share four letters, and nothing else.
# Membership test: the answer is true or false
"go" in languages
# Loop: each item passes through the variable in turn
for language in languages:
print(language)The rule of thumb is easy to hold on to: look at what comes before. Behind an if, or anywhere a true or false value is expected, in tests membership. Between a for and the colon, it declares a loop.
The consequence is a concrete one. A membership test produces a bool value, one that can be stored or combined with other conditions. The line of a for produces nothing at all.
What it looks at, type by type
What exactly is being searched? The question never changes, but where Python goes to ask it depends on the type of the collection. On the left of the table, what is being questioned; on the right, what the answer really covers.
| Collection | What in tests |
|---|---|
| list and tuple | The presence of an item, compared by equality |
| string | The presence of a whole piece, not only of a character |
| dictionary | The presence of a key, never of a value |
| set | The presence of an item, found through its hash |
| range | Membership of the interval, worked out without walking it |
The dictionary row is the one that surprises the most, and it deserves a pause. "price" in article questions the keys, never the values. Searching the other side has to be spelled out, with in article.values().
That choice does a precise job: it tells whether a key exists before anything touches it, which is how a missing key is told apart from a key that is there but holds None. Without that check, a direct lookup would raise a KeyError.
On a range, in walks nothing: Python works out whether the value falls inside the interval and on the right step. 999_999 in range(10_000_000) therefore answers instantly, where the matching list would take a long while and a lot of memory.
The cost of the test
That computing rather than walking leads to the real subject: two in tests written in exactly the same way do not cost the same.
On a list or a tuple, Python compares the items one by one, until it finds a match or until the end. The test therefore slows down as the collection grows, and the worst case is the one where the value is missing: everything has to be walked before concluding. On a set or a dictionary, the answer lands immediately, whatever the size.
blocked = ["a@example.com", "b@example.com"] # list: walked again on every turn
blocked = {"a@example.com", "b@example.com"} # set: answered immediately
for customer in customers:
if customer.email in blocked:
continueHence the habit worth taking: as soon as the same test comes back inside a loop, convert the collection into a set once, before entering it. Only the punctuation of the line changes, while the running time changes order of magnitude.
not in, and the confusion with is
That leaves the opposite case: checking that a value is missing. It is handled with not in, an operator in its own right rather than a negation applied afterwards. email not in blocked reads in the direction of reading, and it is the form to keep.
One last confusion is worth clearing up, the one with is. in looks for a value inside a collection, is asks whether two names point at the same object in memory. The first works on content, the second on identity.
Writing if value is items instead of the membership test raises no error at all. The line is valid, it simply answers false forever, without signalling anything. A condition that never fires gets spotted far later than a program that crashes.
Frequently asked questions
How can a value be checked as missing?
With not in, written as a single block: if email not in blocked:. The form not (email in blocked) works too, but it forces the inner expression to be read before the negation makes sense.
Why does the test fail on a dictionary?
Because a dictionary answers about its keys, not about its values. To search the values, go through in my_dict.values(); for a full pair, through in my_dict.items(). These three spellings answer three different questions.
Can in be used on objects of your own?
Yes, and it is meant to be. A class that defines the magic method __contains__ decides the answer itself, and Python calls it rather than walking anything. Without that method it falls back on the underlying iterable, which costs a full walk on every single test.