Apache Spark interview questionsQuestion 1 of 5
Apache Spark interview question · Question 1 of 5
What is the difference between a transformation and an action in Spark?
Short answer
A transformation such as select, filter or join describes a new DataFrame from an existing one and is lazy: Spark only records it in a plan. An action such as count, collect, show or a write forces Spark to optimise that plan and actually run a job. Laziness lets Spark combine steps and skip unnecessary work before anything executes.
Detailed explanation
Spark builds a logical plan as you chain transformations. Nothing runs yet. When you call an action, Spark optimises the whole plan (for example pushing filters down and pruning columns), splits it into stages and runs it as tasks on the cluster.
| Transformations (lazy) | Actions (trigger a job) |
|---|---|
select, filter, withColumn |
count, collect, show, take |
join, groupBy().agg() |
write...save() / saveAsTable |
distinct, repartition |
foreach, toPandas |
Narrow and wide
- Narrow transformations (
filter,select) let each output partition depend on a single input partition, so no data moves. - Wide transformations (
groupBy, mostjoins,distinct) need data with the same key together, which causes a shuffle and a new stage.
Example
orders = spark.read.parquet("/data/orders") # nothing runs yet
big = orders.filter("amount > 100").select("customer_id", "amount")
by_customer = big.groupBy("customer_id").sum("amount") # still lazy
by_customer.show() # action: a job runs
Common mistakes
- Thinking
readorfilterruns immediately. - Calling several actions on the same expensive DataFrame without caching, so Spark recomputes it each time.
- Using
collect()on large data, pulling everything to the driver.
Progress is saved in this browser only. No account needed.