Menu

ETL and ELT course · Lesson 1 of 8

ETL vs ELT and Modern Data Pipelines

How modern data pipelines are built: ETL vs ELT, batch vs streaming, layered models, orchestration, idempotency, data quality, observability and recovery.

  • Beginner
  • Pillar guide
  • 2 min read
  • Updated Oct 2026
On this page
  1. Design choices
  2. Where transformation happens: ETL or ELT
  3. When it runs: batch or streaming
  4. How it is organised: layers
  5. Reliability practices
  6. Orchestration
  7. Idempotency
  8. Data quality
  9. Observability
  10. Failure handling and recovery
  11. Change data capture
  12. Put it together

A data pipeline moves data from where it is produced to where it is useful, changing it along the way. Modern pipelines share a set of design choices and a set of reliability practices; this guide covers both.

Design choices

Where transformation happens: ETL or ELT

ETL transforms before loading; ELT loads raw data and transforms inside the warehouse or lakehouse. ELT dominates where compute is elastic, but compliance or format constraints can still require ETL steps.

Read: ETL vs ELT · Practise: Deciding factors

When it runs: batch or streaming

Choose by required freshness. Batch is simpler and cheaper; streaming is justified when decisions depend on near-real-time data and brings event time, state and duplicates to handle.

Read: Batch vs streaming

How it is organised: layers

Raw (as received), staging or silver (typed, cleaned, deduplicated), marts or gold (business models). Each layer is rebuildable from the previous one.

Read: Star schema · Data lakes and the lakehouse

Reliability practices

Orchestration

An orchestrator runs tasks in dependency order, on schedule, with retries and alerts, driving each run from its data interval.

Read: Airflow DAGs, scheduling and retries

Idempotency

Every step must be safe to rerun: deterministic input slice, overwrite or merge, atomic publish and guarded side effects.

Read: Idempotency in data pipelines

Data quality

Check freshness, volume, schema, keys and validity at ingestion, after transformation and before publishing; block on critical failures.

Read: Checks, contracts and failure handling

Observability

Record status, duration, volume, freshness and quality per run; alert first on freshness.

Read: Pipeline observability

Failure handling and recovery

Retry transient failures with backoff, quarantine bad data, fail fast on systemic errors, and plan backfills.

Read: Reliability and retry design

Change data capture

Replicate databases from their change logs with ordered, idempotent merges.

Read: CDC patterns and failure modes

Put it together

Work through the scalable batch pipeline case study and build the CSV to warehouse project.

By Data Career Hub Editorial · Last reviewed Oct 2026 · Tool-neutral overview; details and verified examples are in the linked guides

Progress is saved in this browser only. No account needed.

Search
Filter by type