Menu

Course · Data platforms

Delta Lake

Delta Lake adds ACID transactions, schema enforcement and time travel to files in a data lake, which is the foundation of the lakehouse pattern.

Lessons
6
Interview questions
3
Projects & case studies
5
Reading time
~1 h

About this course

Delta Lake is an open table format that stores data as Parquet files plus a transaction log. That log turns a folder of files into a table with atomic commits, consistent reads and a history you can query.

Learn what the transaction log provides first, then schema enforcement and evolution, then maintenance such as compaction.

Your progress

Saved in this browser only

Practise

Course structure

Lessons

Work through the lessons in order. Completed lessons show a tick; lessons you have opened are outlined.

Start here

The complete overview of the course in one read.

  1. Data Lakes, Lakehouse Architecture and Delta LakeFrom data lakes to the lakehouse: file formats, table formats, Delta Lake transactions, schema evolution, medallion layers, maintenance and governance in one guide.Intermediate2 min

Beginner

Core concepts you will use every day.

  1. JSON vs Parquet for Analytics PipelinesWhy raw JSON is convenient to land but expensive to query, how Parquet fixes size, types and scan cost, and a practical pattern for converting between them.Beginner2 min

Intermediate

Patterns used in production pipelines.

  1. Delta Lake: Transactions, Schema Evolution and Time TravelSee how Delta Lake's transaction log gives ACID guarantees on data lake files, how schema enforcement and evolution work, and how time travel and VACUUM interact.Intermediate4 min
  2. Delta Lake vs Traditional Data Lake TablesWhat changes when a folder of Parquet files becomes a Delta table: atomic writes, updates and deletes, schema enforcement, time travel and faster metadata.Intermediate2 min
  3. Parquet vs Avro vs ORC for Data EngineeringCompare Parquet, ORC and Avro by layout, compression, schema evolution and use case, with a measured size comparison and how column pruning works in Spark.Intermediate2 min

Advanced

Performance, internals and edge cases.

  1. Schema Evolution: Safe Design PatternsPatterns for evolving data schemas without breaking consumers: additive changes, expand-and-contract migrations, contracts, versioned events and handling type changes.Advanced2 min

Projects and case studies

Apply what you learned and prepare material to discuss in interviews.

Projects

System design case studies

Resources

Cheat sheets

Related courses

Plan your learning

Search
Filter by type