Menu

Delta Lake course · Lesson 1 of 6

Data Lakes, Lakehouse Architecture and Delta Lake

From data lakes to the lakehouse: file formats, table formats, Delta Lake transactions, schema evolution, medallion layers, maintenance and governance in one guide.

  • Intermediate
  • Pillar guide
  • 2 min read
  • Updated Oct 2026
On this page
  1. 1. Lakes, warehouses and lakehouses
  2. 2. File formats
  3. 3. Table formats and the transaction log
  4. 4. Schema enforcement and evolution
  5. 5. Medallion layers
  6. 6. Maintenance
  7. 7. Governance
  8. Design it

A data lake stores files cheaply in object storage. A lakehouse adds the reliability of a warehouse on top of those files through an open table format such as Delta Lake. This guide connects the pieces.

1. Lakes, warehouses and lakehouses

Lakes are flexible and cheap but lack transactions and enforced schemas; warehouses are reliable but keep data in their own storage; lakehouses keep data in open formats on object storage while adding transactions and governance.

Read: Lake vs warehouse vs lakehouse · Practise: When to choose each

2. File formats

Use columnar Parquet for analytical data, row-based Avro for messages, and treat JSON and CSV as landing formats only.

Read: Parquet vs Avro vs ORC · JSON vs Parquet

3. Table formats and the transaction log

Delta Lake writes Parquet data files plus a _delta_log of commits. That log provides atomic writes, consistent snapshots, optimistic concurrency, time travel and row-level MERGE, UPDATE and DELETE.

Read: Delta Lake transactions, schema evolution and time travel · Delta vs plain Parquet tables · Practise: What Delta Lake solves

4. Schema enforcement and evolution

Reject mismatched writes by default; evolve schemas deliberately, preferring additive changes and expand-and-contract migrations for breaking ones.

Read: Schema evolution patterns · Practise: When schema evolution is safe

5. Medallion layers

Organise tables into bronze (raw), silver (cleaned and conformed) and gold (business models). Each layer reads only the previous one and can be rebuilt.

Read: Databricks workspace, jobs and lakehouse

6. Maintenance

Compact small files, cluster by common filters, and VACUUM old files with a retention period that respects time travel and long readers.

Read: Partitioning, clustering and data layout

7. Governance

A catalog provides names, permissions, masking and lineage across engines.

Read: Unity Catalog and governance

Design it

Put it together in the scalable lakehouse case study, then build the Kafka → Spark → Delta streaming project.

By Data Career Hub Editorial · Last reviewed Oct 2026 · Describes Delta Lake 3.x; comparable capabilities exist in Apache Iceberg and Apache Hudi

Progress is saved in this browser only. No account needed.

Search
Filter by type