ETL and ELT courseLesson 1 of 8
ETL and ELT course · Lesson 1 of 8
ETL vs ELT and Modern Data Pipelines
How modern data pipelines are built: ETL vs ELT, batch vs streaming, layered models, orchestration, idempotency, data quality, observability and recovery.
On this page
A data pipeline moves data from where it is produced to where it is useful, changing it along the way. Modern pipelines share a set of design choices and a set of reliability practices; this guide covers both.
Design choices
Where transformation happens: ETL or ELT
ETL transforms before loading; ELT loads raw data and transforms inside the warehouse or lakehouse. ELT dominates where compute is elastic, but compliance or format constraints can still require ETL steps.
Read: ETL vs ELT · Practise: Deciding factors
When it runs: batch or streaming
Choose by required freshness. Batch is simpler and cheaper; streaming is justified when decisions depend on near-real-time data and brings event time, state and duplicates to handle.
Read: Batch vs streaming
How it is organised: layers
Raw (as received), staging or silver (typed, cleaned, deduplicated), marts or gold (business models). Each layer is rebuildable from the previous one.
Read: Star schema · Data lakes and the lakehouse
Reliability practices
Orchestration
An orchestrator runs tasks in dependency order, on schedule, with retries and alerts, driving each run from its data interval.
Read: Airflow DAGs, scheduling and retries
Idempotency
Every step must be safe to rerun: deterministic input slice, overwrite or merge, atomic publish and guarded side effects.
Read: Idempotency in data pipelines
Data quality
Check freshness, volume, schema, keys and validity at ingestion, after transformation and before publishing; block on critical failures.
Read: Checks, contracts and failure handling
Observability
Record status, duration, volume, freshness and quality per run; alert first on freshness.
Read: Pipeline observability
Failure handling and recovery
Retry transient failures with backoff, quarantine bad data, fail fast on systemic errors, and plan backfills.
Read: Reliability and retry design
Change data capture
Replicate databases from their change logs with ordered, idempotent merges.
Read: CDC patterns and failure modes
Put it together
Work through the scalable batch pipeline case study and build the CSV to warehouse project.
Progress is saved in this browser only. No account needed.