Menu

Course · Distributed processing

Apache Spark

Understand how Spark turns your code into jobs, stages and tasks, and why partitions, shuffles and data skew drive performance.

Lessons
5
Interview questions
9
Projects & case studies
19
Reading time
~1 h

About this course

Spark is a distributed processing engine. Your code builds a plan; Spark splits it into stages at shuffle boundaries and runs each stage as parallel tasks over partitions. Most tuning comes down to controlling how much data moves between executors and how evenly it is spread.

Start with partitions and shuffles, then move to join strategies and Adaptive Query Execution.

Your progress

Saved in this browser only

Practise

Course structure

Lessons

Work through the lessons in order. Completed lessons show a tick; lessons you have opened are outlined.

Start here

The complete overview of the course in one read.

  1. Apache Spark Architecture: Driver, Executors and Cluster ManagersHow a Spark application is built: the driver and executors, standalone, YARN and Kubernetes cluster managers, deploy modes, and SparkSession versus SparkContext.Beginner14 min

Beginner

Core concepts you will use every day.

  1. Spark RDD Fundamentals: Transformations, Actions and Key-Value OperationsLearn Spark's RDD layer: lineage and partitions, transformations versus actions, map versus flatMap, reduceByKey versus groupByKey, mapPartitions and closures.Beginner17 min

Intermediate

Patterns used in production pipelines.

  1. Spark Execution Model: Lazy Evaluation, Jobs, Stages and TasksHow Spark turns lazy DataFrame code into a DAG of jobs, stages and tasks, why shuffles split stages, and what to look for in each tab of the Spark UI.Intermediate15 min

Advanced

Performance, internals and edge cases.

  1. Spark Partitions, Shuffles and Data SkewLearn how Spark splits data into partitions, why shuffles are expensive, how to recognise data skew and which fixes (AQE, broadcast, salting) apply.Advanced4 min
  2. Spark Adaptive Query Execution and OptimizationHow Adaptive Query Execution re-plans Spark queries at run time: coalescing shuffle partitions, switching join strategies and splitting skewed partitions.Advanced2 min

Projects and case studies

Apply what you learned and prepare material to discuss in interviews.

Projects

System design case studies

Resources

Cheat sheets

Related courses

Plan your learning

Search
Filter by type