Menu

Course · Distributed processing

PySpark

PySpark is the Python API for Apache Spark. Learn DataFrames, joins, window functions and how partitions and shuffles decide performance.

Lessons
8
Interview questions
7
Projects & case studies
2
Reading time
~1 h

About this course

PySpark lets you write distributed data processing in Python. DataFrames look familiar if you know SQL or pandas, but they are executed lazily across a cluster, so how data is partitioned and moved decides whether a job takes minutes or hours.

Learn DataFrames and joins first, then window functions, then the execution model (jobs, stages, tasks) and how to read the Spark UI.

Your progress

Saved in this browser only

Practise

Course structure

Lessons

Work through the lessons in order. Completed lessons show a tick; lessons you have opened are outlined.

Start here

The complete overview of the course in one read.

  1. PySpark Fundamentals for Data EngineersThe PySpark fundamentals in one guide: DataFrames and schemas, lazy evaluation, joins, window functions, UDF alternatives, partitions, shuffles and how to debug performance.Beginner2 min

Beginner

Core concepts you will use every day.

  1. PySpark DataFrames, Columns and SchemasBuild PySpark DataFrames with explicit StructType schemas, transform them with select, withColumn and filter, handle nulls, and cast types safely under Spark 4 ANSI mode.Beginner24 min
  2. PySpark Transformations vs ActionsUnderstand lazy evaluation in PySpark: transformations build a plan, actions run it, and narrow versus wide transformations decide where Spark shuffles data.Beginner2 min
  3. PySpark Aggregations, Pivot and Temporary ViewsAggregate PySpark DataFrames with groupBy and agg, build subtotals with rollup and cube, pivot and unpivot data, and share DataFrames with SQL through temp views.Beginner16 min
  4. Reading and Writing Data in PySpark: Parquet, CSV, JSON, ORC, Avro and JDBCRead and write Parquet, CSV, JSON, ORC and Avro in PySpark, capture corrupt records, load JDBC tables in parallel and use partition discovery to prune folders.Beginner25 min

Intermediate

Patterns used in production pipelines.

  1. PySpark Joins and Join StrategyWrite correct PySpark joins (inner, left, anti, semi), avoid duplicate-column and fan-out bugs, and understand when Spark broadcasts or sort-merges a join.Intermediate3 min
  2. PySpark Window Functions: Ranking, Lag and Running TotalsUse PySpark window functions for top-N per group, previous-row comparisons and running totals, and avoid the single-partition trap of windows without partitionBy.Intermediate3 min
  3. PySpark UDFs and Safer AlternativesWhen Python UDFs are slow or risky in PySpark, how built-in functions and pandas UDFs compare, and what Arrow-optimised UDFs in Spark 4 change.Intermediate3 min

Projects and case studies

Apply what you learned and prepare material to discuss in interviews.

Projects

Resources

Cheat sheets

Related courses

Plan your learning

Search
Filter by type