Resources
Data Engineering resources and cheat sheets
Reference material for quick revision. Each sheet says which versions it covers.
Cheat sheets
- Cheat sheetAirflowA quick Airflow reference: TaskFlow DAGs, schedules and data intervals, retries, templating, sensors, trigger rules and the Airflow 3 CLI commands you use most.
- Cheat sheetSpark interviewThe Spark concepts interviewers ask about most, in one page: lazy evaluation, stages and shuffles, joins, partitions, skew, AQE, caching and Spark 4 defaults.
- Cheat sheetSystem designA one-page framework for data engineering system design interviews: requirements, estimates, architecture, storage, processing, reliability, quality and trade-offs.
- Cheat sheetDatabricksA quick Databricks reference: Unity Catalog names and grants, Delta table operations, medallion layers, job design and the compute choices that keep costs down.
- Cheat sheetdbtA quick dbt reference: models, ref and source, materializations, incremental models, tests, snapshots and the commands and selectors used to run a project.
- Cheat sheetKafkaA quick Kafka reference: topics, partitions, keys, replication, consumer groups, offsets, delivery semantics and the CLI commands for inspecting a cluster.
- Cheat sheetPySparkA quick PySpark reference: reading and writing data, column expressions, joins, aggregations, window functions and the settings that matter for performance.
- Cheat sheetPythonA quick Python reference for pipelines: files and CSV, JSON, dates, collections, generators and batching, error handling, logging and testing with pytest.
- Cheat sheetSnowflakeA quick Snowflake reference: warehouses, loading data, time travel, cloning, clustering, query profiling and the habits that keep compute costs under control.
- Cheat sheetSQLA quick SQL reference for Data Engineers: join types, aggregation, window functions, deduplication, upserts and the mistakes that change row counts.