Course · Streaming & orchestration
Airflow
Airflow schedules and orchestrates pipelines as DAGs. Learn scheduling, task dependencies, retries and idempotent task design.
- Lessons
- 4
- Interview questions
- 2
- Projects & case studies
- 3
- Reading time
- ~1 h
About this course
Airflow is a workflow orchestrator: it decides what runs, when, in what order and what happens on failure. It does not process your data itself. It triggers tools that do.
Learn DAG structure and scheduling first, then retries, backfills and how to design tasks that are safe to rerun.
Your progress
Saved in this browser onlyPractise
- InterviewAirflow interview questionsThe full list with difficulty, type and a box to tick off each one.
- Cheat sheetAirflow Cheat SheetA quick Airflow reference: TaskFlow DAGs, schedules and data intervals, retries, templating, sensors, trigger rules and the Airflow 3 CLI commands you use most.
- InterviewAll interview questionsEvery question across all topics in one filterable list.
Course structure
Lessons
Work through the lessons in order. Completed lessons show a tick; lessons you have opened are outlined.
Beginner
Core concepts you will use every day.
- Airflow DAG Fundamentals, TaskFlow and Dynamic DAGsBuild Airflow 3 DAGs from first principles: tasks and dependencies, the TaskFlow API, params and Jinja templating, dynamic DAGs and dynamic task mapping.
- Airflow Operators, Hooks, Providers, Branching and Trigger RulesUse BashOperator and PythonOperator well, write custom operators and hooks, pick provider packages, run pods, and control flow with branching and trigger rules.
Intermediate
Patterns used in production pipelines.
- Airflow Sensors and Deferrable OperatorsWait for files, other DAGs and external jobs without wasting workers: sensor modes, timeouts, ExternalTaskSensor, FileSensor, triggers and what replaced Smart Sensors.
- Airflow DAGs, Scheduling, Retries and Task DependenciesLearn how Airflow DAGs define task order, how scheduling intervals really work, and why retries are only safe when each task is idempotent.
Projects and case studies
Apply what you learned and prepare material to discuss in interviews.
Projects
System design case studies
- AdvancedDesign a Data Observability SystemA data platform runs 3,000 tables across a warehouse and a lakehouse, fed by Airflow, dbt, Spark and streaming jobs. Problems are usually found by business users hours later. Design a data observability system that monitors pipelines and data automatically, detects freshness, volume, schema and distribution anomalies, finds the likely root cause through lineage, and drives incidents to resolution against defined SLAs.
- IntermediateDesign a Batch Ingestion FrameworkA data team writes a new pipeline by hand for every source, and now runs 150 slightly different jobs pulling from databases, SFTP drops, object storage and REST APIs. Design a reusable, metadata-driven batch ingestion framework that onboards a new source through configuration, lands data reliably and idempotently in the lakehouse, and is easy to operate, backfill and monitor.
Resources
Cheat sheets
Related courses
- PythonPython glues pipelines together: ingestion, validation, orchestration and PySpark jobs. Focus on functions, generators, error handling and testable code.
- KafkaKafka is a distributed log used for streaming data. Learn topics, partitions, consumer groups and delivery semantics before building streaming pipelines.
- Apache SparkUnderstand how Spark turns your code into jobs, stages and tasks, and why partitions, shuffles and data skew drive performance.