Menu

Python interview question · Question 3 of 4

What is a generator and why can it help with large datasets?

  • Easy
  • conceptual / coding
  • ~6 min
  • High relevance
  • 2 min read
  • Updated Oct 2026

Short answer

A generator is a function that uses yield to produce values one at a time, pausing between them, instead of building a whole list. Because only the current item is held in memory, you can stream a file, a query result or paginated API data of any size with roughly constant memory. The trade-off is that a generator can be consumed only once and has no length or random access.

On this page
  1. Detailed explanation
  2. Where it helps in data engineering
  3. The single-use trap
  4. Common mistakes

Detailed explanation

import sys

squares_list = [i * i for i in range(1_000_000)]
squares_gen = (i * i for i in range(1_000_000))

print(sys.getsizeof(squares_list) > 1_000_000, sys.getsizeof(squares_gen) < 1_000)
print(sum(squares_gen) == sum(squares_list))
True True
True

The list object alone takes megabytes; the generator object is a few hundred bytes regardless of how many values it will produce, and gives the same sum.

Where it helps in data engineering

  • Reading large files line by line.
  • Iterating over database cursors or paginated APIs.
  • Chaining transformation stages that each handle one record.
  • Feeding a loader in fixed-size batches.

The single-use trap

gen = (x for x in [1, 2, 3])
print(list(gen), list(gen))
[1, 2, 3] []

The second pass is silently empty.

Common mistakes

  1. Converting a generator to a list “to be safe”, which removes the benefit.
  2. Calling len() on it.
  3. Consuming it twice and not noticing.

By Data Career Hub Editorial · Last reviewed Oct 2026 · Examples run on Python 3.12

Progress is saved in this browser only. No account needed.

Search
Filter by type