Definition
Crossing two lists of values, chaining them one after another, or spotting consecutive pairs in a series of readings: these needs come up constantly, each with its own loop. The real cost is the list that each loop builds in memory before it can be used, when the program often only needs one value at a time.
itertools is the standard library module that answers this problem: functions that take an iterable and hand back another one, computing each value on demand rather than building the whole result at once. Nothing to install, it has always been part of Python. Here is what that looks like with combinations:
from itertools import combinations
team = ["Ana", "Bilal", "Chloe", "Dan"]
for pair in combinations(team, 2):
print(pair)
# ('Ana', 'Bilal')
# ('Ana', 'Chloe')
# ('Ana', 'Dan')
# ...Every function therefore returns an iterator: an object empty at creation, that produces its values one at a time. That is what lets it work on sequences longer than the available memory, and it is also the source of the module's most common trap, described further down.
What it actually replaces
Nothing itertools offers is impossible to write by hand: a few lines of loops almost always get the same result. The gain is not a new capability, but faster reading: a named function states what it does, where a nested loop forces the reader to rebuild the intent.
The table below sets each function against the loop it saves writing.
| Function | What it saves writing |
|---|---|
chain(a, b) | One loop per source, one after the other |
product(a, b) | A for nested inside another, plus one per added dimension |
combinations(a, 2) | A double loop over positions, plus a test so the same pair is not counted twice |
islice(stream, 10) | A counter and a manual exit to stop at the tenth item |
pairwise(readings) | Reaching the next item by its position, with the overshoot that comes with it |
The clearest case remains product. Crossing four settings means stacking four nested loops, and adding one more if a fifth shows up. With product, the list of settings becomes a single argument: adding a dimension no longer touches the code, only the input.
The single-use iterator trap
This is the mistake almost everyone discovering the module makes, and it raises no exception at all: the code runs without complaint. An iterator is consumed, no value is kept by default, and walking through it twice means starting from nothing.
Here is what that looks like in practice:
from itertools import chain
everything = chain([1, 2], [3, 4])
print(list(everything)) # [1, 2, 3, 4]
print(list(everything)) # []The visual variant is just as common: printing the object directly gives <itertools.chain object at 0x10f3a2b40> rather than the expected values. This is not a bug, it is the price of producing on demand: nothing is kept except what gets explicitly stored.
The fix comes down to one simple rule: convert to a list with list() as soon as the result needs to be used twice, and leave the iterator as it is otherwise.
groupby does not group what people expect
The name groupby suggests it gathers every occurrence of a given key. That is not the case: it only groups consecutive items sharing a key, nothing more. As soon as another key slips in, grouping starts over from scratch.
Here is what that looks like on unsorted sales data:
from itertools import groupby
sales = [("Paris", 10), ("Lyon", 4), ("Paris", 7)]
for city, rows in groupby(sales, key=lambda s: s[0]):
print(city, [row[1] for row in rows])
# Paris [10]
# Lyon [4]
# Paris [7]Paris shows up twice instead of once, split apart by Lyon. The fix is to sort on the same key before calling groupby, reusing the very same lambda on both sides.
Skipping that sort raises no error: the code runs and returns a result that looks correct, just wrong. That is what makes the mistake so costly to catch once it reaches production.
This is not a design flaw: groupby never holds more than one item in memory, which lets it group a log file of several gigabytes already sorted by date without loading it whole. On unsorted, small data, a dictionary keyed by group does just as well, more simply.
When it is of no use
A module offering this many functions makes it tempting to find a spot for every one of them, which is the wrong reflex. On a sequence walked once, a list comprehension reads better than starmap or filterfalse, and zip with enumerate already cover the everyday case without importing anything.
The memory argument only matters past a certain volume. On a thousand items, the lazy version and the version that builds a full list perform about the same, and the second one reads better. It earns its place on long streams, large files and endless series.
Frequently asked questions
Does it need installing before use?
No, it has been part of the standard library since Python 2.3, and a plain import is enough. Searching for a package of the same name on pip would bring back an abandoned namesake at best, something else entirely at worst.
How does it differ from a generator written by hand?
Not at all in principle: a generator also produces its values on demand. The difference is practical: these functions are written in C, faster than their Python equivalent, and above all already named, tested, and recognised by whoever reads the code six months later.
Why does my loop over count() never stop?
Because count, cycle and repeat produce endless series: no StopIteration will ever come along to stop them. The limit has to be set by hand, with islice for a fixed number of items, with takewhile when the condition depends on the values, or with an explicit exit inside the loop.