Menu

Data Engineering interview question · Question 6 of 6

How would you design a CDC pipeline?

  • Hard
  • architecture / scenario
  • ~12 min
  • High relevance
  • 2 min read
  • Updated Oct 2026

Short answer

I would use log-based change data capture: a connector reads the database's transaction log and publishes one event per row change, with the operation, row values and log position, to Kafka topics keyed by primary key so each row's changes stay in order. A consumer applies micro-batches to the target table with MERGE, keeping only the latest position per key and ignoring events older than what is stored, which makes replays harmless. I also plan the initial snapshot, delete handling, schema changes and monitoring of lag.

Detailed explanation

Walk through the design in this order:

  1. Capture: log-based CDC (reads the transaction log, captures deletes, adds no query load) rather than polling updated_at (misses deletes and intermediate changes).
  2. Transport: Kafka topic per table, keyed by primary key, so ordering holds per row.
  3. Apply: micro-batch MERGE into the target. Deduplicate to the latest log position per key; update only when the incoming position is newer.
  4. Bootstrap: consistent snapshot plus the log position at snapshot time; stream from that position.
  5. Deletes: hard delete, or a soft-delete flag if history is needed.
  6. Schema changes: additive changes flow through; breaking changes pause and alert.
  7. Operations: monitor connector lag, consumer lag and end-to-end latency; keep Kafka retention longer than the longest expected outage.

See the full CDC system design case study.

Common mistakes

  1. Applying events in arrival order without comparing log positions.
  2. Forgetting deletes.
  3. No plan for the initial load.

By Data Career Hub Editorial · Last reviewed Oct 2026

Progress is saved in this browser only. No account needed.

Search
Filter by type