Definition
You have a log file of several gigabytes and you want to count the error lines in it. The reflex is to write a function that reads everything, stores each line in a list, then hands that list back. It works on a small file and brings the machine down on a large one: everything has to fit in memory before the smallest computation. The generator exists to avoid that.
A generator is a function that produces its values one at a time, as they are asked for, instead of computing them all and handing them back in one block. A single keyword gives it away: its body holds a yield where an ordinary function would write return. The caller therefore gets no longer a result, but an object able to supply several of them.
def squares(n):
for i in range(n):
yield i * i # hands a value over, then puts the function on hold
for value in squares(5):
print(value)
# 0 1 4 9 16No list exists during that walk: every square is computed when the loop asks for it, handed over, then forgotten. A generator therefore stays within a few hundred bytes. You still use it like any other collection, because it is an iterator, hence an iterable that for walks through without knowing anything about how it is built.
What the call really does
One question is left hanging by that example: when does the body of the function run? Calling squares(5) runs not a single line of it. The call builds a generator object and stops there. The code starts on the first value requested, pauses on the yield, then resumes at that exact spot on the next request, its variables untouched. Watch when each print comes out.
def counter():
print("starting")
yield 1
print("in between")
yield 2
c = counter() # not one line of the body has run yet
next(c) # prints starting, returns 1
next(c) # resumes after the first yield, prints in between, returns 2
next(c) # nothing left to produce: StopIterationThat suspension is the whole point of the mechanism: an ordinary function runs from start to finish in one go and forgets everything on the way out, a generator keeps its place and its variables. When it has nothing left to produce, it raises StopIteration, the end signal the for loop catches on its own: that is why a walk finishes cleanly without any stop condition.
Generator or list
Both are walked through the same way, and nothing, when reading a loop, tells you which one you are handling. What separates them is the memory, the moment the values are computed and the number of readings they allow. The table below puts them side by side.
| Criterion | List | Generator |
|---|---|---|
| Memory used | Every value | One at a time |
| When values are computed | At construction | On every request |
| Number of walks | Unlimited | One only |
| len and slicing | Available | Unavailable |
| Endless sequence | Impossible | Natural |
The gap this chart shows is not a nuance: at one million values, the list takes tens of thousands of times more room. The choice is therefore quickly made. A generator as soon as the source is large and handled as it flows, a list as soon as counting, sorting or rereading is needed.
The generator stays at its fixed weight while the list grows with the data. That is where it becomes the only workable option, on a file too large to load at once.
The orders of magnitude are those of CPython on a 64-bit machine: about 28 bytes per integer plus an 8-byte pointer inside the list, and about 200 bytes for the generator object, whatever sequence it produces. Measure your own with sys.getsizeof.
The most frequent confusion is about range, taken for a generator because it stores nothing. It is a cousin, not a member of the family: it can be reread as many times as you like and accepts len as well as slicing.
The generator expression
Writing a whole function with yield for a simple sequence of values is often out of proportion. Python offers something shorter: a list comprehension placed between parentheses rather than square brackets becomes a generator expression. The result is no longer a list built in one shot, but a generator.
heavy = [i * i for i in range(10000000)] # ten million values in memory
light = (i * i for i in range(10000000)) # nothing at all, until it is walked
total = sum(i * i for i in range(10000000)) # the parentheses of sum are enoughThe second line computes nothing: it prepares a recipe, not a result. The third one deserves a stop, because when the expression is the only argument of a call, the parentheses of the call are enough. This is the most common shape in daily work, well ahead of functions holding a yield: one square bracket swapped for a parenthesis defuses a computation that would fill up memory.
The single-walk trap
A generator wears out. Once it reaches the end it does not rewind: the second reading hands nothing back and, above all, no error warns you.
gen = (i for i in range(3))
print(list(gen)) # [0, 1, 2] the generator is now empty
print(list(gen)) # [] no error, just nothing leftThe classic symptom shows up when the same generator goes through two treatments: the first consumes everything and works, the second works on emptiness and returns zero. No exception, no message, just a wrong result that nothing distinguishes from a right one. Two ways out: convert to a list, or build a fresh generator.
A generator also refuses certain operations. len(gen) raises a TypeError, reading by position does too, and learning how many items there are means consuming them all: what has not been built yet cannot be counted.
A generator that reads a file opens nothing until it is walked. Returning it from a function where the file closes on the way out gives a valid object that fails on the first value, with a ValueError about a closed file. Consume it before the closing.
Frequently asked questions
What is the difference between yield and return?
Both hand a value back, but not at the same price. return ends the function: the value leaves, the local state disappears, nothing will resume. yield delivers a value then puts the function on hold, variables included. A bare return stays legal inside a generator: it hands nothing back and stops the production.
Is a generator always faster?
No, and that is a widespread misunderstanding. On a complete walk, a list is often slightly faster, because producing every value on demand carries a small overhead. The real gain lies in memory and in the delay before the first value, decisive when the walk stops early or the sequence is very long.
How can a list be obtained from a generator?
With list(my_generator), which consumes everything and puts the values back in memory. The operation cancels the very saving that was sought and is only justified when several walks are needed; on an endless sequence, it never finishes. The Python course revisits that trade-off through the handling of large files.