Menu

Airflow course · Lesson 4 of 5

Airflow Scheduling: Intervals, Logical Dates, Catchup, Timetables and Assets

How the Airflow 3 scheduler creates runs: cron and presets, logical dates and data intervals, catchup pitfalls, timetables, asset-aware scheduling and deadline alerts.

  • Intermediate
  • 26 min read
  • Updated Oct 2026
On this page
  1. Running the examples
  2. Scheduler internals
  3. What it is
  4. How it works
  5. Pitfalls
  6. In interviews
  7. Schedule intervals and cron
  8. What it is
  9. How it works
  10. Pitfalls
  11. In interviews
  12. Execution date vs logical date
  13. What it is
  14. How it works
  15. Pitfalls
  16. In interviews
  17. Catchup and start_date pitfalls
  18. What it is
  19. How it works
  20. Pitfalls
  21. In interviews
  22. Timetables
  23. What it is
  24. How it works
  25. Pitfalls
  26. In interviews
  27. Data-aware scheduling with assets
  28. What it is
  29. How it works
  30. Pitfalls
  31. In interviews
  32. SLAs and deadline alerts
  33. What it was
  34. Why it went
  35. What replaces it: deadline alerts
  36. Pitfalls
  37. In interviews
  38. Retries and idempotency
  39. Practice questions
  40. Key takeaways

Most Airflow confusion is about time: why a run labelled 1 March started on 2 March, why turning a DAG on created 300 runs, or why a DAG that “should have run” never did. This lesson explains how the scheduler decides when to create runs, what the logical date and data interval mean (and how Airflow 3 changed them), and the alternatives to time-based scheduling: timetables and assets. It ends with deadline alerts, the Airflow 3 replacement for SLAs.

Running the examples

The examples use Airflow 3.3 with a metadata database created by airflow db migrate. run_dag registers an in-memory DAG and runs it once with dag.test(), as in the earlier lessons. next_runs asks a timetable which runs it would create, using the same timetable classes the scheduler uses.

import contextlib, io, tempfile, warnings
from datetime import datetime, timedelta
import pendulum
from airflow.dag_processing.bundles.manager import DagBundlesManager
from airflow.dag_processing.dagbag import DagBag, sync_bag_to_db
from airflow.models.xcom import XComModel
from airflow.timetables.base import TimeRestriction
from airflow.utils.session import create_session

def run_dag(dag, **test_kwargs):
    """Register an in-memory DAG, run it once with dag.test(), return (run state, task states, return values)."""
    with contextlib.redirect_stdout(io.StringIO()):
        DagBundlesManager().sync_bundles_to_db()
        bag = DagBag(dag_folder=tempfile.mkdtemp())
        bag.dags[dag.dag_id] = dag
        sync_bag_to_db(bag, "dags-folder", None)
        dr = dag.test(**test_kwargs)
    states, values = {}, {}
    with create_session() as session:
        for ti in dr.get_task_instances(session=session):
            states[ti.task_id] = str(ti.state)
        for x in session.query(XComModel).filter_by(dag_id=dag.dag_id, run_id=dr.run_id, key="return_value"):
            values[x.task_id] = XComModel.deserialize_value(x)
    return str(dr.state), states, values

def next_runs(timetable, start, n=3):
    """The first n runs a timetable would create from start, as (logical date, data interval, run after)."""
    last, out = None, []
    for _ in range(n):
        info = timetable.next_dagrun_info(
            last_automated_data_interval=last,
            restriction=TimeRestriction(earliest=start, latest=None, catchup=True),
        )
        if info is None:
            break
        fmt = lambda d: d.strftime("%m-%d %H:%M")
        out.append(f"logical {fmt(info.logical_date)} | interval {fmt(info.data_interval.start)} -> "
                   f"{fmt(info.data_interval.end)} | runs after {fmt(info.run_after)}")
        last = info.data_interval
    return out

Scheduler internals

What it is

The scheduler is the component that turns DAG definitions into work. It decides when a DAG run should exist, which task instances are ready, and hands ready tasks to an executor, which runs them on workers. It does not run your task code.

How it works

An Airflow 3 deployment has several cooperating processes, all sharing the metadata database:

Component Job
DAG processor (airflow dag-processor) Parses DAG files from DAG bundles, serialises DAGs into the database, records import errors. In Airflow 3 it is always a separate process from the scheduler.
Scheduler (airflow scheduler) Creates DAG runs from timetables and asset events, sets task instances to scheduled and queued, enforces concurrency, sends work to the executor.
Executor (runs inside the scheduler) Local, Celery, Kubernetes, or several at once. Starts each queued task somewhere.
Workers Run task code. In Airflow 3 they talk to the API server through the Task Execution API rather than connecting to the metadata database directly.
Triggerer (airflow triggerer) Runs async triggers for deferred tasks.
API server (airflow api-server) Serves the UI, the REST API (/api/v2) and the Task Execution API.

Each scheduler loop (simplified):

  1. Create DAG runs for DAGs whose next run is due (run_after has passed), up to max_active_runs per DAG, and for DAGs whose asset conditions are met.
  2. Examine running DAG runs: evaluate each task instance’s dependencies (upstream states, trigger rules, depends_on_past) and move ready ones to scheduled.
  3. Critical section: pick scheduled task instances in priority order while respecting [core] parallelism, pool slots, max_active_tasks per DAG and per-task limits, and move them to queued.
  4. Executor heartbeat: send queued tasks to the executor and collect state changes.
  5. Housekeeping: time out stuck queued tasks, detect tasks whose heartbeats stopped, update the next run dates.

You can run more than one scheduler for high availability and throughput. They coordinate through row-level locks in the database (SELECT ... FOR UPDATE SKIP LOCKED), so the database must support that: PostgreSQL, or MySQL 8. SQLite supports only one scheduler and is for local development.

Pitfalls

  • “My DAG doesn’t appear” is usually a DAG processor problem (import error, file not in the bundle, parse timeout), not a scheduler problem. Check import errors first.
  • Slow parsing delays scheduling, because the scheduler works from what the DAG processor has serialised.
  • A busy metadata database slows every component. It is the most common bottleneck at scale.
  • Tasks stuck in queued usually mean the executor has no free capacity (workers down, pool full, Kubernetes quota), not that the scheduler is broken.

In interviews

“Walk me through what happens between writing a DAG file and a task running.” Cover parsing and serialisation by the DAG processor, DAG run creation from the timetable, dependency checks, the critical section with pools and concurrency limits, the executor, and the worker. Mentioning the Airflow 3 split (separate DAG processor, workers going through the API server) and HA schedulers with row locks shows depth.

Schedule intervals and cron

What it is

The schedule argument tells the scheduler when to create runs. Airflow 2.4 merged the old schedule_interval and timetable arguments into schedule, and Airflow 3 removed the old names.

How it works

schedule accepts several kinds of value, and Airflow turns each into a timetable:

from airflow.sdk import DAG, Asset
from airflow.providers.standard.operators.empty import EmptyOperator

for schedule in ["@daily", "0 6 * * *", timedelta(hours=4), None, "@once", "@continuous", [Asset("s3://lake/orders/")]]:
    extra = {"max_active_runs": 1} if schedule == "@continuous" else {}
    d = DAG(f"schedule_{abs(hash(str(schedule)))}", schedule=schedule, start_date=datetime(2026, 1, 1), **extra)
    label = "[Asset('s3://lake/orders/')]" if isinstance(schedule, list) else str(schedule)
    print(f"{label:30} -> {type(d.timetable).__name__}")
@daily                         -> CronTriggerTimetable
0 6 * * *                      -> CronTriggerTimetable
4:00:00                        -> DeltaTriggerTimetable
None                           -> NullTimetable
@once                          -> OnceTimetable
@continuous                    -> ContinuousTimetable
[Asset('s3://lake/orders/')]   -> AssetTriggeredTimetable
Value Meaning
Cron string, e.g. "0 6 * * *" minute hour day-of-month month day-of-week; here 06:00 every day
Presets @hourly, @daily, @weekly, @monthly, @quarterly, @yearly (cron shortcuts at midnight / the start of the period)
timedelta a fixed frequency measured from the previous run
None never scheduled; manual, API or asset-driven triggers only
"@once" one run, then never again
"@continuous" start the next run as soon as the previous one finishes (requires max_active_runs=1)
asset or asset expression run when upstream data is updated (see data-aware scheduling)
a timetable object full control, see Timetables

Cron times are interpreted in the timezone of the DAG’s start_date (UTC unless you pass a timezone-aware date; [core] default_timezone is utc by default). With a non-UTC timezone, cron follows local wall-clock time across daylight-saving changes.

Pitfalls

  • Reading cron fields in the wrong order; 0 6 * * 1 is 06:00 on Mondays, not “every 6 minutes”.
  • Timezone-naive start_date while the business thinks in local time: the job runs an hour off for half the year.
  • timedelta(days=1) and @daily are not the same: the timedelta is relative to the previous run, cron is aligned to the clock.
  • Using schedule_interval= in Airflow 3 code; it is no longer accepted.

In interviews

Be able to read and write cron expressions quickly (“every 15 minutes on weekdays between 8 and 18” is */15 8-18 * * 1-5), list the presets and explain that None means externally triggered only. Mention timezone-aware start dates.

Execution date vs logical date

What it is

Every scheduled DAG run has a logical date: the timestamp that identifies which period of data it is responsible for. Airflow 1 and early Airflow 2 called it the execution date, a name that misled almost everyone into thinking it was when the run executes. Airflow 2.2 introduced logical_date and data intervals, and Airflow 3 removed execution_date (and related context names such as next_execution_date, prev_execution_date, tomorrow_ds, yesterday_ds) entirely.

How it works

A run also has a data interval (data_interval_start, data_interval_end) and a run_after time, the earliest moment the scheduler may start it. How these relate depends on the timetable, and Airflow 3 changed the default for cron schedules:

  • Airflow 2 default, CronDataIntervalTimetable: a daily run covers the previous day. The run with logical date 1 March has the interval 1 March 00:00 to 2 March 00:00 and starts after the interval ends, on 2 March. This is the famous “the run labelled 1 March runs on 2 March”.
  • Airflow 3 default, CronTriggerTimetable: a cron schedule simply fires at the cron time. The run that fires at 1 March 06:00 has logical date 1 March 06:00 and an empty data interval (start = end = 1 March 06:00).
from airflow.timetables.interval import CronDataIntervalTimetable
from airflow.timetables.trigger import CronTriggerTimetable

start = pendulum.datetime(2026, 3, 1, tz="UTC")
print("Airflow 3 default: CronTriggerTimetable")
print(*next_runs(CronTriggerTimetable("0 6 * * *", timezone="UTC"), start), sep="\n")
print("Airflow 2 default: CronDataIntervalTimetable")
print(*next_runs(CronDataIntervalTimetable("0 6 * * *", timezone="UTC"), start), sep="\n")
print("CronTriggerTimetable with interval=1 day")
print(*next_runs(CronTriggerTimetable("0 6 * * *", timezone="UTC", interval=timedelta(days=1)), start), sep="\n")
Airflow 3 default: CronTriggerTimetable
logical 03-01 06:00 | interval 03-01 06:00 -> 03-01 06:00 | runs after 03-01 06:00
logical 03-02 06:00 | interval 03-02 06:00 -> 03-02 06:00 | runs after 03-02 06:00
logical 03-03 06:00 | interval 03-03 06:00 -> 03-03 06:00 | runs after 03-03 06:00
Airflow 2 default: CronDataIntervalTimetable
logical 03-01 06:00 | interval 03-01 06:00 -> 03-02 06:00 | runs after 03-02 06:00
logical 03-02 06:00 | interval 03-02 06:00 -> 03-03 06:00 | runs after 03-03 06:00
logical 03-03 06:00 | interval 03-03 06:00 -> 03-04 06:00 | runs after 03-04 06:00
CronTriggerTimetable with interval=1 day
logical 02-28 06:00 | interval 02-28 06:00 -> 03-01 06:00 | runs after 03-01 06:00
logical 03-01 06:00 | interval 03-01 06:00 -> 03-02 06:00 | runs after 03-02 06:00
logical 03-02 06:00 | interval 03-02 06:00 -> 03-03 06:00 | runs after 03-03 06:00

Read the three blocks carefully:

  1. Airflow 3’s default fires at 06:00 each day; logical date and run time are the same.
  2. The Airflow 2 behaviour labels each run with the start of the day it processes and runs it a day later.
  3. CronTriggerTimetable(..., interval=...) keeps “fire at the cron time” but gives the run a data interval ending at the fire time (here the previous 24 hours), with the logical date at the interval start.

If your tasks rely on data_interval_start/data_interval_end (or on ds meaning “the day being processed”), either set [scheduler] create_cron_data_intervals = True to get the Airflow 2 behaviour for every cron string, or pass the timetable explicitly on the DAGs that need it. Timedelta schedules have the matching switch [scheduler] create_delta_data_intervals. Both are False by default in Airflow 3.

Manual and API-triggered runs can have no logical date at all in Airflow 3 (logical_date=None when you trigger without one), so code that assumes it is always set should handle None.

Pitfalls

  • Migrating a DAG from Airflow 2 to 3 without checking data intervals: a query using ds suddenly processes “today” instead of “yesterday”. Decide explicitly.
  • Using datetime.now() instead of the run’s dates: backfills and reruns process the wrong period.
  • Assuming logical_date is when the run started. Use dag_run.start_date for that.
  • Templates with {{ execution_date }} fail in Airflow 3 with an undefined-variable error.

In interviews

“What is the difference between execution date and logical date?” A strong answer: they are the same concept renamed, because “execution date” sounded like the run time; the logical date identifies the data period, the data interval gives its bounds, and run_after says when it may run. Then add the Airflow 3 change: cron strings use CronTriggerTimetable by default, so data intervals are empty unless you opt in, and execution_date is removed.

Catchup and start_date pitfalls

What it is

When a DAG is unpaused, the scheduler compares its start_date with now. With catchup on, it creates a run for every interval between start_date (or the last run) and now. With catchup off, it creates only the most recent one and moves on. Airflow 3 changed the default [scheduler] catchup_by_default to False; in Airflow 2 it was True.

How it works

  • catchup=True with start_date a year in the past and a daily schedule: about 365 runs are created, limited by max_active_runs (16 by default) at a time.
  • catchup=False: only the latest interval is scheduled. Historical periods are run deliberately with a backfill (airflow backfill create, the UI or the REST API), covered in the operations lesson.
  • start_date is not “the date of the first run”. With data intervals, the first run’s interval starts at start_date and runs after it ends; with trigger timetables, the first run fires at the first cron time on or after start_date.
  • Changing start_date on an existing DAG does not rewrite history, and moving it earlier does not create the missing runs unless catchup is on.
  • end_date stops scheduling after a date.

Pitfalls

  • Dynamic start dates: start_date=datetime.now() (or days_ago(1), removed in Airflow 3) moves every time the file is parsed, so the next run is always in the future and never created.
  • Accidental floods: turning on a DAG with catchup=True and an old start_date floods workers and external systems. Set catchup explicitly and cap max_active_runs.
  • Runs that depend on the previous run (depends_on_past=True) plus catchup: one failure blocks every later run.
  • Not idempotent tasks plus catchup or backfill: duplicates. Overwrite or merge per interval.
  • Pausing a DAG for a week with catchup off: the missed week is silently skipped. Backfill it if the data matters.

In interviews

Classic questions are “Why didn’t my DAG run?” (dynamic start date, paused DAG, start date in the future, the run is not due until the interval ends) and “I turned on a DAG and it created hundreds of runs. Why?” (catchup with an old start date). Mention that Airflow 3’s default is catchup=False and that you should set it explicitly either way.

Timetables

What it is

A timetable is the object behind every schedule. It decides each run’s logical date, data interval and run_after. When cron and timedeltas cannot express your calendar (business days only, different times on weekends, a fixed list of dates), you pass a timetable instead of a string.

How it works

Built-in timetables in Airflow 3 (import from airflow.sdk, or airflow.timetables.* in Airflow 2):

Timetable Use
CronTriggerTimetable(cron, timezone=..., interval=...) Fire at cron times; optional data interval ending at the fire time. Airflow 3 default for cron strings.
CronDataIntervalTimetable(cron, timezone) Airflow 2 style: one run per interval between cron ticks, run after the interval.
DeltaTriggerTimetable(delta) / DeltaDataIntervalTimetable(delta) The same two styles for fixed frequencies.
MultipleCronTriggerTimetable(*crons, timezone=...) Several cron expressions in one schedule.
EventsTimetable(event_dates=[...]) An explicit list of dates, such as quarter-ends or a release calendar.
AssetOrTimeSchedule(timetable=..., assets=...) Run on a time schedule and when assets update.
from airflow.timetables.events import EventsTimetable
from airflow.timetables.trigger import MultipleCronTriggerTimetable

weekday_and_weekend = MultipleCronTriggerTimetable("0 6 * * 1-5", "0 9 * * 6,0", timezone="UTC")
print(*next_runs(weekday_and_weekend, pendulum.datetime(2026, 3, 6, tz="UTC"), n=4), sep="\n")

quarter_ends = EventsTimetable(event_dates=[pendulum.datetime(2026, 3, 31, tz="UTC"),
                                            pendulum.datetime(2026, 6, 30, tz="UTC")])
print(*next_runs(quarter_ends, pendulum.datetime(2026, 1, 1, tz="UTC")), sep="\n")

london = CronTriggerTimetable("0 2 * * *", timezone="Europe/London")      # across the clock change
print(*next_runs(london, pendulum.datetime(2026, 3, 28, tz="Europe/London")), sep="\n")
logical 03-06 06:00 | interval 03-06 06:00 -> 03-06 06:00 | runs after 03-06 06:00
logical 03-07 09:00 | interval 03-07 09:00 -> 03-07 09:00 | runs after 03-07 09:00
logical 03-08 09:00 | interval 03-08 09:00 -> 03-08 09:00 | runs after 03-08 09:00
logical 03-09 06:00 | interval 03-09 06:00 -> 03-09 06:00 | runs after 03-09 06:00
logical 03-31 00:00 | interval 03-31 00:00 -> 03-31 00:00 | runs after 03-31 00:00
logical 06-30 00:00 | interval 06-30 00:00 -> 06-30 00:00 | runs after 06-30 00:00
logical 03-28 02:00 | interval 03-28 02:00 -> 03-28 02:00 | runs after 03-28 02:00
logical 03-29 01:00 | interval 03-29 01:00 -> 03-29 01:00 | runs after 03-29 01:00
logical 03-30 01:00 | interval 03-30 01:00 -> 03-30 01:00 | runs after 03-30 01:00

Friday 6 March ran at 06:00, the weekend at 09:00, Monday at 06:00 again; the events timetable stopped after its last date; and the London schedule kept 02:00 local time: the times are printed in UTC, so the run moves from 02:00 to 01:00 UTC when British Summer Time begins on 29 March.

A custom timetable subclasses airflow.timetables.base.Timetable and implements next_dagrun_info() (the next scheduled run) and infer_manual_data_interval() (the interval for a manual run), plus serialize/deserialize if it has parameters. It must be registered through a plugin so the scheduler can load it:

from airflow.plugins_manager import AirflowPlugin
from airflow.timetables.base import DagRunInfo, DataInterval, Timetable

class BusinessDaysTimetable(Timetable):
    def infer_manual_data_interval(self, *, run_after):
        ...  # return DataInterval(start, end) for a manual run

    def next_dagrun_info(self, *, last_automated_data_interval, restriction):
        ...  # return DagRunInfo.interval(start, end), skipping weekends and holidays, or None

class BusinessDaysPlugin(AirflowPlugin):
    name = "business_days_timetable"
    timetables = [BusinessDaysTimetable]

Pitfalls

  • Writing a custom timetable when MultipleCronTriggerTimetable or EventsTimetable would do. Custom timetables need a plugin, and they run in the scheduler, so bugs affect scheduling.
  • Slow or I/O-heavy logic in a custom timetable. It is called inside the scheduler loop.
  • Forgetting to restart components after changing a plugin.

In interviews

“How would you run a DAG only on business days, at different times at weekends, or on a list of dates?” Name the built-ins first, and describe a custom timetable (two methods, registered as a plugin) for holiday calendars. Explaining that cron strings are just shorthand for a timetable shows you understand the model.

Data-aware scheduling with assets

What it is

Data-aware scheduling runs a DAG when the data it depends on has been updated, instead of at a fixed time. A producer task declares that it updates an asset (a logical reference to data, identified by a URI or name); a consumer DAG is scheduled on that asset. Airflow 2.4 introduced this as Datasets; Airflow 3 renamed them Assets (Dataset became Asset, DatasetAlias became AssetAlias, and the UI, API and context names changed to match).

How it works

  • Producer: put the asset in a task’s outlets. When the task succeeds, Airflow records an asset event. The task can attach metadata to the event (outlet_events[asset].extra = {...}, or yield Metadata(asset, extra)).
  • Consumer: schedule=[asset_a, asset_b] (all must update), or an expression with & (and) and | (or). The run is created once the condition is met; it can read the triggering events from triggering_asset_events.
  • Airflow does not look at the data itself. An asset event means “the producer task succeeded”, nothing more.
  • @asset defines an asset and the DAG that materialises it in one decorator, and AssetAlias lets a task decide at run time which asset it updated. AssetWatcher (with an event-driven trigger, for example a message queue) can update an asset from outside Airflow.
from airflow.sdk import Asset, asset, dag, task

orders = Asset("s3://lake/orders/", extra={"owner": "sales"})
customers = Asset(uri="s3://lake/customers/", name="customers")

@dag(schedule="@daily", start_date=datetime(2026, 1, 1), catchup=False)
def orders_producer():
    @task(outlets=[orders])
    def write_orders(*, outlet_events, ds=None):
        # ... write s3://lake/orders/dt=<ds>/ ...
        outlet_events[orders].extra = {"rows": 120, "partition": ds}
        return "written"
    write_orders()

with DAG("needs_both", schedule=(orders & customers), start_date=datetime(2026, 1, 1)) as needs_both:
    EmptyOperator(task_id="build_mart")
with DAG("needs_either", schedule=(orders | customers), start_date=datetime(2026, 1, 1)) as needs_either:
    EmptyOperator(task_id="refresh_cache")

for d in (needs_both, needs_either):
    print(d.dag_id, type(d.timetable).__name__, type(d.timetable.asset_condition).__name__)

print(run_dag(orders_producer(), logical_date=pendulum.datetime(2026, 3, 1, tz="UTC"))[:2])

from airflow.models.asset import AssetEvent
with create_session() as session:
    for event in session.query(AssetEvent).filter(AssetEvent.source_dag_id == "orders_producer"):
        print("asset event:", event.asset.uri, event.extra, "from", f"{event.source_dag_id}.{event.source_task_id}")

@asset(schedule="@daily", uri="s3://lake/daily_summary/")
def daily_summary():
    return "summary built"          # the DAG 'daily_summary' materialises this asset

print(type(daily_summary).__name__, daily_summary.name, daily_summary.uri)
needs_both AssetTriggeredTimetable AssetAll
needs_either AssetTriggeredTimetable AssetAny
('success', {'write_orders': 'success'})
asset event: s3://lake/orders {'rows': 120, 'partition': '2026-03-01'} from orders_producer.write_orders
AssetDefinition daily_summary s3://lake/daily_summary

The producer run recorded an asset event with its metadata. In a running deployment, the scheduler would now check needs_either (condition met, run created) and needs_both (still waiting for customers).

Pitfalls

  • Treating an asset event as a data-quality guarantee. It only means a task with that outlet succeeded; add checks in the producer.
  • URIs are identifiers and must match exactly (Airflow normalises some forms, such as a trailing slash, as the uri in the output shows). Define assets once in a shared module and import them.
  • Several producer updates before the consumer runs are combined into one consumer run, not queued one per event. Read triggering_asset_events if you need each one.
  • Mixing time and data triggers by hand. Use AssetOrTimeSchedule for “daily, or earlier if the data arrives”.
  • Using Dataset imports in Airflow 3 code; use Asset from airflow.sdk.

In interviews

“How do you trigger a DAG when another DAG’s output is ready?” The modern answer is assets: producer outlets, consumer schedule=[asset], &/| conditions, event metadata. Compare with ExternalTaskSensor (polling, needs date alignment) and TriggerDagRunOperator (push, couples the producer to the consumer). Say “Datasets were renamed Assets in Airflow 3”.

SLAs and deadline alerts

What it was

An SLA (service level agreement) in Airflow 2 was a per-task sla=timedelta(...) measured from the start of the DAG run’s data interval end. When a task had not finished in time, the scheduler recorded an SLA miss, sent an email and called the DAG’s sla_miss_callback. It was widely used for “the finance tables must be ready by 07:00” alerts.

Why it went

The legacy SLA feature was removed in Airflow 3.0. It was evaluated inside the scheduler loop (which cost scheduler time on every DAG), it fired only for scheduled runs, it measured from a confusing reference point, and misses for tasks that never started were easy to miss. In Airflow 3.x passing sla= to a task only logs a warning and the value is ignored:

with warnings.catch_warnings(record=True) as caught:
    warnings.simplefilter("always")
    with DAG("sla_demo", schedule=None, start_date=datetime(2026, 1, 1)) as sla_dag:
        EmptyOperator(task_id="load", sla=timedelta(hours=1))
print([str(w.message) for w in caught if "SLA" in str(w.message)])
print("task still has an sla attribute:", hasattr(sla_dag.get_task("load"), "sla"))
['The SLA feature is removed in Airflow 3.0, replaced with Deadline Alerts in >=3.1']
task still has an sla attribute: False

What replaces it: deadline alerts

Deadline alerts (added in Airflow 3.1) attach a deadline to a DAG run: a reference time, an interval after (or before) it, and a callback to run if the DAG run has not finished by then. Built-in references include DeadlineReference.DAGRUN_QUEUED_AT (time since the run was queued), DAGRUN_LOGICAL_DATE (time since the run’s logical date), FIXED_DATETIME(...) and AVERAGE_RUNTIME(...). Callbacks are an AsyncCallback (runs in the triggerer) or, from Airflow 3.2, a SyncCallback (runs on an executor); either can wrap a notifier (for example a Slack notifier from its provider) or your own function given by import path.

from airflow.sdk import AsyncCallback, DeadlineAlert, DeadlineReference

async def page_on_call(**kwargs):
    ...   # call your paging or chat tool here; must be importable by the triggerer

with DAG(
    "finance_tables",
    schedule="0 5 * * *",
    start_date=datetime(2026, 1, 1),
    catchup=False,
    deadline=DeadlineAlert(
        reference=DeadlineReference.DAGRUN_LOGICAL_DATE,   # measured from 05:00
        interval=timedelta(hours=2),                       # must be done by 07:00
        callback=AsyncCallback(page_on_call, kwargs={"team": "finance-data"}),
    ),
) as finance:
    EmptyOperator(task_id="build_tables")

alert = finance.deadline[0]
print(type(alert).__name__, alert.interval, alert.callback.path)   # the callback is stored by import path
DeadlineAlert 2:00:00 __main__.page_on_call

The deadline is checked by the scheduler, and the callback runs whether or not any task started, which fixes the biggest gap of the old SLAs. For a limit on a single task, use execution_timeout (fail the task) rather than an alert.

Pitfalls

  • Expecting sla= to do anything in Airflow 3. Migrate to deadline alerts.
  • Choosing the wrong reference: DAGRUN_QUEUED_AT catches slow runs, DAGRUN_LOGICAL_DATE catches “not ready by the business deadline” even when the run started late.
  • Callbacks that are not importable where they run, or that block (async callbacks run on the triggerer’s event loop).
  • Using dagrun_timeout when you wanted an alert: it fails the run instead of notifying.

In interviews

Interviewers still ask “How do you implement SLAs in Airflow?”. Explain the Airflow 2 sla parameter and sla_miss_callback, say it was removed in Airflow 3.0, and describe deadline alerts (reference, interval, callback) as the replacement, plus execution_timeout and dagrun_timeout for hard limits and failure callbacks for alerts.

Retries and idempotency

Scheduling decides when a run happens; retries, reruns and backfills decide how often the same interval is processed. Every one of them is only safe if each task is idempotent: rerunning it for the same interval leaves the same final state. Overwrite the run’s partition (or delete-then-insert in one transaction), upsert on a unique key, or write to a temporary location and swap it in atomically; never append blindly. The Python tutorial on an idempotent loader shows the pattern in code, and the operations lesson covers retries, backfills and clearing in depth.

Practice questions

A @daily DAG in Airflow 2 shows a run for 1 March that started on 2 March. Why? Would Airflow 3 do the same?

In Airflow 2, cron schedules used CronDataIntervalTimetable: the run for logical date 1 March covers 1 March 00:00 to 2 March 00:00 and can only start once that interval has ended, so it starts on 2 March. In Airflow 3 cron strings use CronTriggerTimetable by default, so the run fires at midnight on 1 March with logical date 1 March and an empty data interval, unless you set create_cron_data_intervals = True or pass a data-interval timetable.

You deployed a DAG with start_date=datetime.now() and it never runs. Why?

The start date is re-evaluated every time the DAG file is parsed, so it keeps moving forward and the first interval never becomes due. Use a fixed, timezone-aware start date.

What happens when you unpause a daily DAG whose start_date is 1 January, on 1 October?

With catchup=True, the scheduler creates a run for every missed day (about 270), up to max_active_runs at a time. With catchup=False (the Airflow 3 default) it creates only the latest run; earlier days must be backfilled deliberately.

What replaced execution_date in Airflow 3, and what else changed with it?

logical_date (in templates {{ logical_date }}, in code dag_run.logical_date), together with data_interval_start/data_interval_end. Related names such as next_execution_date, prev_execution_date, yesterday_ds and tomorrow_ds were removed. Manual runs may have no logical date.

The marketing DAG must run when both the orders and customers tables have been refreshed, whichever order they finish in. How?

Declare orders and customers as assets in the producer tasks’ outlets and schedule the marketing DAG with schedule=(orders & customers) (or a list of both). The scheduler creates a run once both have new events since the last run. Assets were called datasets in Airflow 2.

How do you alert if the finance DAG has not finished by 07:00, in Airflow 3?

Use a deadline alert: DeadlineAlert(reference=DeadlineReference.DAGRUN_LOGICAL_DATE, interval=...) measured from the scheduled time (or FIXED_DATETIME), with an async or sync callback that notifies the team. The old sla parameter and sla_miss_callback were removed in Airflow 3.0.

Why can running two schedulers on SQLite not work?

Multiple schedulers coordinate through row-level locking (SELECT ... FOR UPDATE SKIP LOCKED), which SQLite does not support. HA schedulers need PostgreSQL or MySQL 8; SQLite is only for local development with a single scheduler.

Key takeaways

  • The DAG processor parses and serialises DAGs; the scheduler creates runs, checks dependencies and queues tasks within concurrency limits; the executor starts them.
  • schedule takes cron strings, presets, timedeltas, None, @once, @continuous, assets or timetables; each becomes a timetable.
  • logical_date replaced execution_date. In Airflow 3 cron strings fire at the cron time with empty data intervals unless you opt into data intervals.
  • Set a fixed, timezone-aware start_date and an explicit catchup (the Airflow 3 default is False); backfill history on purpose.
  • Use built-in timetables (multiple crons, event dates) before writing a custom one, which needs a plugin.
  • Assets (formerly datasets) schedule consumers on producer success; SLAs were removed in 3.0 and deadline alerts replace them.

By Data Career Hub Editorial · Last reviewed Oct 2026 · Examples run on Apache Airflow 3.3.2 (Task SDK 1.3.2) with Python 3.11 and a SQLite metadata database. Timetable outputs come from the scheduler-side timetable classes; Airflow 2 behaviour is described where it differs. The scheduler and HA configuration is described from the documentation, not executed.

Progress is saved in this browser only. No account needed.

Search
Filter by type