Apache Spark interview questionsQuestion 2 of 5
Apache Spark interview question · Question 2 of 5
Explain Spark jobs, stages and tasks.
Short answer
When you call an action, Spark creates a job to compute it. The job is divided into stages at shuffle boundaries, because each wide transformation needs all data for a key before continuing. Each stage runs as a set of tasks, one per partition, executed in parallel on executor cores. With Adaptive Query Execution enabled, one action may show up as several jobs, because Spark re-plans after each shuffle stage.
Detailed explanation
| Concept | Created by | Count determined by |
|---|---|---|
| Job | An action (count, write, show) |
One or more per action |
| Stage | Shuffle boundaries in the plan | Number of wide transformations + 1 (roughly) |
| Task | A partition in a stage | Number of partitions in that stage |
Example
Reading a dataset with 8 input partitions, filtering, then groupBy("country").count() and show():
- Job: triggered by
show(). - Stage 1: read + filter + partial aggregation, 8 tasks (one per input partition), ending with shuffle write.
- Stage 2: shuffle read + final aggregation, as many tasks as shuffle partitions (often fewer after AQE coalescing).
Driver versus executors
The driver plans and schedules; executors run tasks and store shuffle and cached data. A slow driver usually means too much collect() or too many tiny tasks to schedule.
Common mistakes
- Saying each transformation is a stage (only shuffles create new stages).
- Equating tasks with executors (tasks map to partitions; executors provide cores).
Progress is saved in this browser only. No account needed.