- Explain the iteration protocol (
iter,next,StopIteration) and describe what happens inside aforloop - Write an iterator class with
__iter__and__next__, and a generator function withyield - Use lazy evaluation to build infinite sequences and processing pipelines
- Apply
yield fromandsend()
A for loop works with lists, strings, dictionaries, files, range and zip — even with a 10 GB log file that would never fit in memory. How can one loop walk through such different objects? The answer is the iteration protocol, a tiny agreement between for and the objects it loops over. Once you understand it, you can write objects that produce values only when needed, process endless data streams and build memory-friendly processing pipelines. The main tool for this is one of Python's most powerful features: generators.
Iterables and iterators
An iterable is any object that can give you an iterator: it has an __iter__() method (lists, strings, dictionaries, sets, files, range). An iterator is the object that actually hands out the values one by one: its __next__() method returns the next value and raises the StopIteration exception when there are no more. The built-in functions iter(x) and next(it) simply call these two methods.
Let's do by hand what for does for us. iter() asks the list for an iterator, and each next() moves it one step forward. When the values run out, the iterator does not return a special value — it raises **StopIteration**:
colors = ['red', 'green', 'blue']
it = iter(colors)
print(type(it).__name__)
print(next(it))
print(next(it))
print(next(it))
try:
next(it)
except StopIteration:
print('StopIteration: the iterator is exhausted')▸ Expected output
list_iterator red green blue StopIteration: the iterator is exhausted
So a for loop is really a while loop that keeps calling next() until StopIteration appears. Below is a working imitation of for. It handles a string and a dictionary in exactly the same way, because both are iterable (looping over a dictionary gives its keys):
def my_for(iterable, action):
it = iter(iterable)
while True:
try:
item = next(it)
except StopIteration:
break
action(item)
my_for('abc', print)
my_for({'x': 1, 'y': 2}, print)▸ Expected output
a b c x y
Two details matter in practice. First, an iterator is itself iterable: its __iter__() returns self, which is why you can write for x in it. Second, an iterator is single-use: it only moves forward and cannot be rewound. A list hands out a new iterator every time you ask, but zip, map, filter, files and generators are iterators themselves:
pairs = zip(['Aysel', 'Murad'], [91, 78])
print(list(pairs))
print(list(pairs))
nums = [1, 2, 3]
print(iter(nums) is iter(nums))
it = iter(nums)
print(iter(it) is it)▸ Expected output
[('Aysel', 91), ('Murad', 78)]
[]
False
TrueYour own iterator class
Any class with __iter__ and __next__ methods works in a for loop, in list() and sum(), in unpacking — in short, everywhere an iterable is accepted. Here is a countdown:
class Countdown:
def __init__(self, start):
self.current = start
def __iter__(self):
return self
def __next__(self):
if self.current <= 0:
raise StopIteration
value = self.current
self.current -= 1
return value
for n in Countdown(3):
print(n)
print(list(Countdown(5)))▸ Expected output
3 2 1 [5, 4, 3, 2, 1]
It works, but it takes a lot of ceremony for a simple idea: we had to keep the state in self.current by hand and raise StopIteration ourselves. Generators do the same job with far less code.
Generators: functions that pause
A function that contains the keyword **yield is a generator function. Calling it does not run its body — it returns a generator object**, which is an iterator. Each next() runs the body up to the next yield, hands out the value and freezes the function together with all its local variables. The following next() continues exactly where it stopped. When the function ends, the generator raises StopIteration automatically.
def countdown(start):
print('start')
while start > 0:
yield start
start -= 1
print('done')
gen = countdown(2)
print(type(gen).__name__)
print(next(gen))
print(next(gen))
print(next(gen, 'no more values'))▸ Expected output
generator start 2 1 done no more values
start is printed only at the first next(), not when countdown(2) is called. next(gen, default) returns the default instead of raising StopIteration.Compare this with the Countdown class: the same behaviour in five lines, and the state (start) is just a local variable. Because nothing runs when a generator function is called, errors inside it also appear only at the first next().
Lazy evaluation: values on demand
Generators are lazy: they calculate a value only when someone asks for it. A generator expression — a comprehension in round brackets — is the lazy twin of a list comprehension. The list below stores a million numbers (about 8 MB on a 64-bit computer), while the generator object takes only about 200 bytes, however long the sequence is. But remember: it is single-use.
import sys
squares_list = [n * n for n in range(1_000_000)]
squares_gen = (n * n for n in range(1_000_000))
print(sys.getsizeof(squares_list) > 1_000_000)
print(sys.getsizeof(squares_gen) < 500)
print(sum(squares_gen))
print(sum(squares_gen))▸ Expected output
True True 333332833333500000 0
| Property | List comprehension [...] | Generator expression (...) |
|---|---|---|
| When it is computed | all at once | each item on request |
| Memory | all items | one item at a time |
| Looping again | as often as you like | only once |
len() and indexing | yes | no |
| Infinite sequence | impossible | possible |
Laziness makes two powerful patterns possible. An infinite generator is perfectly fine as long as you take only as many values as you need (itertools.islice takes the first n). A chain of generators forms a processing pipeline: each stage pulls one item at a time from the previous one, so even a file of many gigabytes is processed line by line with almost constant memory:
from itertools import islice
def naturals():
n = 1
while True:
yield n
n += 1
evens = (n for n in naturals() if n % 2 == 0)
print(list(islice(evens, 5)))
with open('server.log', 'w', encoding='utf-8') as f:
f.write('INFO start\nERROR disk full\nINFO ok\nERROR timeout\nWARN slow\n')
with open('server.log', encoding='utf-8') as f:
lines = (line.rstrip('\n') for line in f)
errors = (line for line in lines if line.startswith('ERROR'))
messages = (line.split(' ', 1)[1] for line in errors)
for msg in messages:
print(msg)▸ Expected output
[2, 4, 6, 8, 10] disk full timeout
yield from and two-way generators
yield from iterable passes on every value of another iterable, including another generator. It is the natural tool for recursion, for example to flatten nested lists of any depth:
def flatten(items):
for item in items:
if isinstance(item, list):
yield from flatten(item)
else:
yield item
data = [1, [2, 3, [4, 5]], [], [[6]], 7]
print(list(flatten(data)))▸ Expected output
[1, 2, 3, 4, 5, 6, 7]
yield from is also an expression: its value is whatever the sub-generator returns with return. And yield itself is an expression — generator.send(value) resumes the generator and makes the yield expression evaluate to that value. So a generator can receive data, not only produce it. Before the first send(), the generator must be advanced to its first yield with next():
def numbers(values):
total = 0
for v in values:
yield v
total += v
return total
def pipeline():
subtotal = yield from numbers([1, 2, 3])
print('subtotal:', subtotal)
yield from 'ab'
print(list(pipeline()))▸ Expected output
subtotal: 6 [1, 2, 3, 'a', 'b']
def running_average():
total, count = 0, 0
average = None
while True:
value = yield average
total += value
count += 1
average = total / count
avg = running_average()
next(avg)
print(avg.send(10))
print(avg.send(20))
print(avg.send(60))▸ Expected output
10.0 15.0 30.0
total, count) between calls — no class was needed.Write a generator chunks(items, size) that splits a sequence into pieces of length size and yields them one by one (the last piece may be shorter). Thanks to slicing, it should work with both a list and a string.
def chunks(items, size):
# yield slices of `size` items; the last one may be shorter
...
print(list(chunks([1, 2, 3, 4, 5, 6, 7], 3)))
print(list(chunks('abcde', 2)))▸ Expected output
[[1, 2, 3], [4, 5, 6], [7]] ['ab', 'cd', 'e']
Write a generator fibonacci() that yields the Fibonacci numbers (0, 1, 1, 2, 3, 5, …) forever, and print the first 10 of them with islice.
from itertools import islice
def fibonacci():
# yield 0, 1, 1, 2, 3, 5, ... forever
...
print(list(islice(fibonacci(), 10)))▸ Expected output
[0, 1, 1, 2, 3, 5, 8, 13, 21, 34]
Key points
forcallsiter()once, thennext()untilStopIterationappears.- An iterable has
__iter__; an iterator also has__next__and is single-use. - A function with
yieldreturns a generator when called; its body runs lazily, pausing at eachyield. - A generator expression
( ... )uses constant memory; a list comprehension[ ... ]builds the whole list. yield fromdelegates to another iterable and evaluates to the sub-generator'sreturnvalue.
Check yourself
10 questions. Every correct answer earns XP.
next(it) on an exhausted iterator?