Python interview questionsQuestion 3 of 4
Python interview question · Question 3 of 4
What is a generator and why can it help with large datasets?
Short answer
A generator is a function that uses yield to produce values one at a time, pausing between them, instead of building a whole list. Because only the current item is held in memory, you can stream a file, a query result or paginated API data of any size with roughly constant memory. The trade-off is that a generator can be consumed only once and has no length or random access.
On this page
Detailed explanation
import sys
squares_list = [i * i for i in range(1_000_000)]
squares_gen = (i * i for i in range(1_000_000))
print(sys.getsizeof(squares_list) > 1_000_000, sys.getsizeof(squares_gen) < 1_000)
print(sum(squares_gen) == sum(squares_list))
True True
True
The list object alone takes megabytes; the generator object is a few hundred bytes regardless of how many values it will produce, and gives the same sum.
Where it helps in data engineering
- Reading large files line by line.
- Iterating over database cursors or paginated APIs.
- Chaining transformation stages that each handle one record.
- Feeding a loader in fixed-size batches.
The single-use trap
gen = (x for x in [1, 2, 3])
print(list(gen), list(gen))
[1, 2, 3] []
The second pass is silently empty.
Common mistakes
- Converting a generator to a list “to be safe”, which removes the benefit.
- Calling
len()on it. - Consuming it twice and not noticing.
Progress is saved in this browser only. No account needed.