Menu

Intermediate project · Project 4 of 8

Large-Scale Batch Processing Pipeline

Process a large public dataset (several gigabytes or more) with PySpark into partitioned, query-ready tables, and document how you found and fixed the main performance bottleneck.

  • Intermediate
  • PySpark · Parquet or Delta Lake · Local Spark or a small cloud cluster
  • 2 min read
  • Updated Oct 2026

Requirements

  • Read a dataset of at least several gigabytes
  • Clean, join with a lookup table and aggregate
  • Write partitioned Parquet or Delta output by date
  • Rerun any date without duplicating output
  • Record a before-and-after performance analysis from the Spark UI

Technology stack

PySpark, Parquet or Delta Lake, Local Spark or a small cloud cluster

Dataset

Use a large public dataset with a clear licence, such as public trip records or open web-analytics samples. Record the source and licence in your README.

Business context

Interviewers for Spark roles want evidence that you have worked with data too big for a laptop’s memory and know how to find bottlenecks. This project produces exactly that evidence: a working job and a written performance analysis.

Architecture

  1. Raw files in a raw/ folder or bucket.
  2. PySpark job parameterised by processing date.
  3. Broadcast join with a small lookup table.
  4. Aggregation to the reporting grain.
  5. Partitioned output with overwrite for the processed date.
A single rerunnable batch job; each date is processed independently.

Write the performance analysis as you go: screenshots of the stage timeline, the task-duration spread, shuffle sizes and what changed after your fix. Report the timings you actually observed on your hardware.

By Data Career Hub Editorial · Last reviewed Oct 2026

Progress is saved in this browser only. No account needed.

Search
Filter by type