Menu

Data Engineering interview question · Question 4 of 6

How would you investigate a suddenly slower data pipeline?

  • Medium
  • debugging / scenario
  • ~10 min
  • High relevance
  • 2 min read
  • Updated Oct 2026

Short answer

First I confirm and scope it: which run, which step, since when, and whether the output is still correct. Then I compare a slow run with a recent good one: input volume, the execution plan, stage and task timings, and anything that changed (code deploys, configuration, upstream schemas, cluster size). The usual causes are data growth, a new skewed key, small files, a changed join strategy, resource contention or an upstream delay, and I fix the specific cause rather than just adding compute.

Detailed explanation

1. Scope

  • Is it every run or one? Since which date? One task or the whole DAG?
  • Is the output still correct? (A slowdown caused by a join explosion also produces wrong data.)

2. Compare with a good run

Compare Where
Input rows and bytes Source metrics, audit tables
Stage durations and task-time distribution Spark UI / engine query profile
Physical plan (join strategies, shuffles) SQL tab, EXPLAIN
Recent changes Deploy history, config, upstream schema
Resources Cluster size, queueing, concurrency

3. Common causes and fixes

  • Volume growth → partition pruning, incremental processing, scaling.
  • New skew → AQE skew join, isolate hot keys, salting.
  • Small files → compaction.
  • Plan change (broadcast no longer chosen) → statistics, hints, threshold review.
  • Upstream delay → it is waiting, not slow; fix the dependency or alerting.

4. Prevent recurrence

Track duration and input volume per run, alert on deviation from the recent baseline, and record root causes.

Common mistakes

  1. Adding nodes before understanding the cause.
  2. Looking only at total duration instead of per-step and per-task timings.

By Data Career Hub Editorial · Last reviewed Oct 2026

Progress is saved in this browser only. No account needed.

Search
Filter by type