Menu

Cheat sheet · Sheet 2 of 10

Apache Spark Interview Cheat Sheet

The Spark concepts interviewers ask about most, in one page: lazy evaluation, stages and shuffles, joins, partitions, skew, AQE, caching and Spark 4 defaults.

  • 2 min read
  • Updated Oct 2026
On this page
  1. Execution model
  2. Narrow vs wide
  3. Joins
  4. Partitions
  5. Skew
  6. AQE (on by default)
  7. Caching
  8. Spark 4 defaults worth knowing
  9. Classic answers in one breath

Execution model

Term One-line answer
Transformation Lazy; describes a new DataFrame
Action Triggers execution (count, collect, write)
Job Created by an action (AQE may run one action as several jobs)
Stage Set of tasks between shuffle boundaries
Task Work on one partition, run by one core
Driver Plans and schedules; receives collect() results
Executor Runs tasks, holds cache and shuffle data

Narrow vs wide

Narrow: select, filter, withColumn, union. Wide (shuffle): groupBy, join (non-broadcast), distinct, repartition, orderBy, windows with partitionBy.

Joins

Strategy When
Broadcast hash One side ≤ autoBroadcastJoinThreshold (10 MB) or hinted
Sort-merge Both sides large
Shuffle hash Some cases via AQE or hints

left_anti = rows with no match; left_semi = rows with a match, no right columns.

Partitions

  • Input: from file sizes and spark.sql.files.maxPartitionBytes (128 MB).
  • Shuffle: spark.sql.shuffle.partitions (200), coalesced by AQE.
  • repartition = full shuffle, up or down; coalesce = merge down, no full shuffle.

Skew

Symptom: a few tasks much slower (max ≫ median in the Spark UI). Fixes: AQE skew join → broadcast small side → filter/isolate hot or NULL keys → salting.

AQE (on by default)

Coalesces shuffle partitions, switches to broadcast joins at run time, splits skewed join partitions. Does not fix aggregation skew or excessive scanning.

Caching

cache() only when a DataFrame is reused by several actions; unpersist() afterwards.

Spark 4 defaults worth knowing

Setting Default
spark.sql.ansi.enabled true (errors on invalid casts and overflow)
spark.sql.adaptive.enabled true
spark.sql.execution.pythonUDF.arrow.enabled true (when PyArrow is available)

Classic answers in one breath

  • Why is my job slow? Check the Spark UI: which stage, task-time spread, shuffle size, spill; then fix skew, shuffles or scanning.
  • Why avoid Python UDFs? Serialisation and an opaque function the optimiser cannot reason about; prefer built-ins or pandas UDFs.
  • Why did a join multiply rows? Non-unique key on the right side.

By Data Career Hub Editorial · Last reviewed Oct 2026 · Defaults checked on Spark 4.2

Search
Filter by type