Course · Data platforms
ETL and ELT
ETL transforms data before loading it; ELT loads first and transforms inside the warehouse or lakehouse. Learn when each fits.
- Lessons
- 8
- Interview questions
- 1
- Projects & case studies
- 4
- Reading time
- ~1 h
About this course
ETL and ELT describe where transformation happens. In ETL it happens in a separate processing step before data reaches the target. In ELT raw data is loaded first and transformed inside the warehouse or lakehouse with SQL.
Neither is universally better. The choice depends on data volume, where compute is cheapest, governance needs and team skills.
Your progress
Saved in this browser onlyPractise
Course structure
Lessons
Work through the lessons in order. Completed lessons show a tick; lessons you have opened are outlined.
Start here
The complete overview of the course in one read.
Beginner
Core concepts you will use every day.
Intermediate
Patterns used in production pipelines.
- Batch vs Streaming Data PipelinesWhen to process data in batches and when to stream it: latency needs, complexity, correctness with late data, cost, and the micro-batch middle ground.
- Idempotency in Data PipelinesMake every pipeline step safe to rerun: deterministic inputs, overwrite and merge patterns, atomic publishing, idempotent consumers and guarded side effects.
- Data Quality: Checks, Contracts and Failure HandlingBuild a data quality practice: which checks to run where, how contracts set expectations between teams, and how to decide whether a failure blocks or alerts.
- Data Pipeline Observability FundamentalsWhat to measure and alert on in data pipelines: run status, duration, volume, freshness, quality results, lag and lineage, so problems are found before consumers notice.
- Data Pipeline Reliability and Retry DesignDesign pipelines that recover on their own: classify failures, retry with backoff and limits, isolate bad records, make steps idempotent and plan backfills.
Projects and case studies
Apply what you learned and prepare material to discuss in interviews.
System design case studies
- AdvancedDesign a Data Observability SystemA data platform runs 3,000 tables across a warehouse and a lakehouse, fed by Airflow, dbt, Spark and streaming jobs. Problems are usually found by business users hours later. Design a data observability system that monitors pipelines and data automatically, detects freshness, volume, schema and distribution anomalies, finds the likely root cause through lineage, and drives incidents to resolution against defined SLAs.
- AdvancedDesign a Data Quality FrameworkBad data keeps reaching dashboards and models: duplicated orders after a retry, a source that silently sent half its rows, negative prices after an upstream change. Design a company-wide data quality framework that lets teams declare expectations on their datasets, enforces them at the right points in batch and streaming pipelines, stops bad data from being published, and routes problems to the people who can fix them.
- AdvancedDesign a Financial Reconciliation PipelineDesign a daily batch pipeline that reconciles the company's internal payment ledger with settlement files from payment service providers (PSPs) and statements from banks, so finance can prove every transaction was received, settled and paid out, and can investigate every difference.
- AdvancedDesign a Near-Zero Downtime Data Platform MigrationDesign the migration of a live on-premises data warehouse, the ETL jobs that load it and the dashboards that read it to a cloud warehouse or lakehouse, so that consumers see no more than a few minutes of disruption and every number can be proved to match before the old system is switched off.
Resources
Cheat sheets
Related courses
- Data modelingData modeling and warehousing: star schemas and grain, fact and dimension design, SCDs, Data Vault and other methods, dbt, semantic layers and incremental models.
- SQLSQL is the core language of data work: querying, transforming and modelling data in warehouses, lakehouses and Spark. Start here before any other tool.
- AirflowAirflow schedules and orchestrates pipelines as DAGs. Learn scheduling, task dependencies, retries and idempotent task design.