Menu

ETL and ELT course · Lesson 7 of 8

Data Pipeline Reliability and Retry Design

Design pipelines that recover on their own: classify failures, retry with backoff and limits, isolate bad records, make steps idempotent and plan backfills.

  • Intermediate
  • 2 min read
  • Updated Oct 2026
On this page
  1. Classify failures
  2. Retry design
  3. Timeouts
  4. Isolate failures
  5. Recovery and backfills
  6. Common mistakes
  7. Key takeaway

Reliable pipelines are not pipelines that never fail. They are pipelines where failures are expected, contained and recoverable, usually without a human.

Classify failures

Class Examples Strategy
Transient Timeouts, throttling, lock contention, preempted nodes Retry with exponential backoff and a cap
Data A malformed record, an unexpected value Quarantine and count; fail if the rate is high
Upstream not ready Source file late Wait (sensor or schedule offset) with a timeout, then alert
Systemic Bad credentials, schema change, bug Fail fast and alert; retrying will not help

Retrying a systemic failure only delays the alert.

Retry design

  • Exponential backoff with jitter: wait longer after each failure, with randomness so many tasks do not retry at the same moment.
  • Limits: a small number of attempts and a maximum delay.
  • Retry the smallest unit: small tasks repeat less work when they fail.
  • Idempotency first: a retry must be safe to run (see idempotency).

Timeouts

A task with no timeout can hang forever and block downstream work silently. Set execution timeouts on tasks and on calls to external services.

Isolate failures

  • Quarantine bad records rather than failing a whole batch for one row.
  • Separate independent sources into separate tasks so one failing source does not stop the others.
  • Use circuit breakers for unstable APIs: after repeated failures, stop calling for a while.

Recovery and backfills

Plan how to recover before you need to: reprocess a date range with the same code path (backfill), limit concurrency so backfills do not starve daily runs, and keep raw data so you can rebuild.

Common mistakes

  1. Unlimited retries without backoff.
  2. Retrying non-transient errors.
  3. No timeouts.
  4. Large tasks that redo hours of work on retry.

Key takeaway

Classify failures, retry only transient ones with backoff and limits, set timeouts, isolate bad data, make every step idempotent and plan backfills in advance.

By Data Career Hub Editorial · Last reviewed Oct 2026 · Tool-neutral; retry options shown in Airflow terms

Progress is saved in this browser only. No account needed.

Search
Filter by type