Airflow interview questionsQuestion 1 of 1
Airflow interview question · Question 1 of 1
How should Airflow retries and idempotency work together?
Short answer
Retries make Airflow rerun a failed task, possibly after it already wrote part of its output, so retries are only safe if the task is idempotent. Each task should process exactly its run's data interval, taken from the run context rather than the current time, and write by overwriting that interval's partition or upserting on a unique key inside a transaction. Then a retry, a manual clear or a backfill all produce the same final state.
Detailed explanation
| Rerun cause | Why idempotency matters |
|---|---|
| Automatic retry | The first attempt may have written partial output |
| Manual “clear” of a task | Someone reruns after a fix |
| Backfill | Old intervals are processed again |
Design rules
- Deterministic input: use the run’s data interval (
data_interval_start,data_interval_end). - Replace, not append: partition overwrite,
MERGE, or build-and-swap. - Atomic writes: transactions or table formats with atomic commits.
- Guard side effects: record that a notification was sent for this run before sending again.
- Retry the right failures: transient errors with backoff; let schema errors fail.
Example settings
from datetime import timedelta
default_args = {
"retries": 3,
"retry_delay": timedelta(minutes=2),
"retry_exponential_backoff": True,
"max_retry_delay": timedelta(minutes=30),
}
Common mistakes
- Append-only loads with retries enabled.
- Using the current time instead of the data interval.
- Very high retry counts that hide a real failure for hours.
Progress is saved in this browser only. No account needed.