Menu

Course · Languages & query

Python

Python glues pipelines together: ingestion, validation, orchestration and PySpark jobs. Focus on functions, generators, error handling and testable code.

Lessons
5
Interview questions
4
Projects & case studies
3
Reading time
~1 h

About this course

In Data Engineering, Python is mostly about reliable plumbing: reading from sources, validating records, loading targets, scheduling work and calling libraries such as PySpark. The difference between a script and a pipeline is error handling, logging and the ability to run the same job twice safely.

Start with functions and modules, then iterators and generators, then exceptions and logging.

Your progress

Saved in this browser only

Practise

Course structure

Lessons

Work through the lessons in order. Completed lessons show a tick; lessons you have opened are outlined.

Start here

The complete overview of the course in one read.

  1. Python for Data EngineeringThe Python a Data Engineer actually uses: structuring pipeline code, data structures, generators and batching, error handling, idempotent loads, testing and orchestration.Beginner2 min

Beginner

Core concepts you will use every day.

  1. Python Functions, Modules and Reusable Data Pipeline CodeStructure Python pipeline code as small pure functions, clear modules and a thin entry point so it is testable, reusable and safe to run from any scheduler.Beginner3 min
  2. Python Data Structures for Data Engineering InterviewsChoose between list, tuple, set, dict, deque and Counter by their operations and costs, with the patterns that come up in data engineering coding interviews.Beginner3 min

Intermediate

Patterns used in production pipelines.

  1. Python Iterators, Generators and Memory-Efficient ProcessingUse Python iterators and generators to process files and API pages lazily, in batches, with flat memory use, and avoid the mistakes that silently exhaust them.Intermediate4 min
  2. Build an Idempotent CSV Loader in PythonBuild a small Python loader that reads CSV files, skips and counts bad rows, and loads into a database so running it twice gives the same result.IntermediateTutorial6 min

Projects and case studies

Apply what you learned and prepare material to discuss in interviews.

Projects

Resources

Cheat sheets

Related courses

Plan your learning

Search
Filter by type