Definition
Walking through a list with a for loop feels completely natural, but someone has to remember how far the walk has gone, otherwise every loop would restart at the first item. That someone is the iterator: an object handing out a source's items one at a time.
Its vocabulary holds to two calls: iter() returns the iterator itself, and next() supplies the following item while moving the cursor forward by one. That tiny contract is enough to walk through a three-item list just as well as a multi-gigabyte file that would never fit in memory.
numbers = [10, 20, 30]
cursor = iter(numbers)
next(cursor) # 10
next(cursor) # 20
next(cursor) # 30The cursor image explains why the object holds no values of its own: it simply moves inside a source that exists elsewhere. Once at the end, it raises StopIteration, the agreed signal for the end of the run.
An iterable and an iterator are not the same thing
The two words look alike, which explains a good share of the behaviour judged strange while learning the language. An iterable is a source: a fresh cursor can be requested from it as many times as wanted. The iterator is that cursor, and it serves only once.
The table below sums up the difference where it matters most.
| Question | Iterable | Iterator |
|---|---|---|
| Common examples | List, tuple, string, dictionary | Result of iter(), generator, open file |
Answers next() | No | Yes |
| Walked twice | Yes, indefinitely | No, once only |
| Knows its size | Often, through len | Almost never |
A list builds a fresh cursor on every pass, hence its endless rereading. A generator or an open file is its own cursor: once emptied, it stays that way, and nothing rewinds it.
What a for loop really does
The for loop looks as if it magically knows how to walk through anything, but it invents nothing: it asks for a cursor, calls next() for as long as it gets a value, and leaves at the end signal. The lines below are the exact translation of for item in collection:.
cursor = iter(collection)
while True:
try:
item = next(cursor)
except StopIteration: # the end signal, not an error
break
handle(item)That equivalence has a direct consequence: any object respecting the protocol is walked with the very same writing, whether the loop faces a list held in memory, a file sitting on disk or a network stream arriving drop by drop.
Taking hold of the cursor
Nothing forces the loop to do everything. Since the cursor keeps its own position, it can just as well be grabbed by hand, have a few items consumed apart, and then hand the rest over to the usual walk, which will pick up exactly where it was left. Common case: the header line of a data file, not to be handled like the others.
with open("sales.csv") as source:
cursor = iter(source)
header = next(cursor) # the first line, set aside
for line in cursor:
handle(line)It is the memory of position that makes this split possible: the same object passes from one hand to the other without losing any ground already covered. An ordinary list cannot do this without explicit slicing or a counter kept up to date.
The trap of the already-consumed object
Code that works on the first read and gives back nothing on the second, without the slightest error: that is the most frequent trap around iterators. It comes from an easy mistake, since Python now hands back iterators where older tutorials show lists, with map, filter, zip, enumerate and every generator.
pairs = zip([1, 2, 3], "abc")
print(list(pairs)) # [(1, 'a'), (2, 'b'), (3, 'c')]
print(list(pairs)) # [] : the cursor is already emptyThe defect never shows up where it was made: the first read succeeds, often far from the object it created, and the second reports an empty collection without warning. One habit settles the matter: materialise with list() ahead of time, or rebuild the object on every pass if the volume forbids keeping it all in memory.
An iterator stays true in a condition even when empty, having neither a length nor a truth value of its own: if results: passes even when nothing is left to read, a normally reliable check turned misleading here.
Writing your own
Understanding the protocol makes it possible to implement it by hand. An object becomes an iterator as soon as it exposes __iter__, which returns the object itself, and __next__, which supplies the next value or raises the end signal. Each of them is a magic method, called by Python, never by the code making use of it.
class Counter:
def __init__(self, end):
self.value = 0
self.end = end
def __iter__(self):
return self # the class is its own cursor
def __next__(self):
if self.value >= self.end:
raise StopIteration
self.value += 1
return self.valueThis class is then walked like any other collection. In practice, a yield inside a function produces the same result in three lines: that is why the generator has almost replaced the hand-written class.
Frequently asked questions
Why is my list of results empty on the second pass?
Because the object being walked was an iterator, not a collection. The first pass emptied it, and the second finds nothing left to read, with no error to flag it. A list() placed on the result at the moment of production settles the problem for good.
How can the number of remaining items be known?
No way exists to know it without consuming the object, and len() fails on it. The only answer is to materialise it into a list, loading all of its content into memory, exactly what the iterator was trying to avoid.
Should an iterator be preferred over a list?
As soon as the data is large or read only once, yes: the memory taken stays constant whatever the volume. For a few dozen items read again and again, the list remains simpler and faster.