Course · Cloud
AWS
AWS for Data Engineers: S3, IAM, Lambda, Glue, Athena, Redshift, EMR, Kinesis, Step Functions, Lake Formation and how they fit into one data platform.
- Lessons
- 8
- Interview questions
- 0
- Projects & case studies
- 2
- Reading time
- ~2 h
About this course
Amazon Web Services is the most common cloud in Data Engineer job descriptions. An AWS data platform is usually a lake in Amazon S3, described by the AWS Glue Data Catalog, loaded by streaming and batch ingestion, transformed with Spark on Glue or EMR, queried with Athena or Redshift, orchestrated with Step Functions or Airflow, governed with IAM and Lake Formation, and watched with CloudWatch.
Start with the map of the AWS data stack, then S3 and IAM, because every other service depends on them. After that, follow the lessons in order or jump to the service your team uses.
Your progress
Saved in this browser onlyPractise
Course structure
Lessons
Work through the lessons in order. Completed lessons show a tick; lessons you have opened are outlined.
Start here
The complete overview of the course in one read.
Beginner
Core concepts you will use every day.
- Amazon S3 for Data EngineersS3 for data lakes: buckets and keys, storage classes, lifecycle rules, versioning, partitioned layouts, events, encryption, access points and multipart upload.
- AWS IAM for Data Engineers: Roles, Policies and Cross-Account AccessIAM for data pipelines: users, groups and roles, identity and resource policies, least privilege, STS, cross-account access, boundaries and policy evaluation.
Intermediate
Patterns used in production pipelines.
- AWS Lambda for Data PipelinesUse AWS Lambda in data pipelines: triggers, S3 and Kinesis processing, concurrency, cold starts, layers, memory and timeout tuning, and error handling with DLQs.
- AWS Glue: Data Catalog, Crawlers and Spark ETL JobsAWS Glue for Data Engineers: the Data Catalog, crawlers, partitions, Spark ETL jobs, DynamicFrames, Glue Studio, job bookmarks, triggers, workflows and cost tuning.
- Amazon Athena: Serverless SQL on S3Query S3 with Amazon Athena: external tables, CTAS, partition projection, performance tuning, workgroups, federated queries and what drives the cost per query.
- Amazon Redshift for Data EngineersRedshift for Data Engineers: architecture, RA3, distribution and sort keys, COPY and UNLOAD, vacuum, WLM, concurrency scaling, Spectrum and Serverless.
Projects and case studies
Apply what you learned and prepare material to discuss in interviews.
Projects
System design case studies
Resources
Related courses
- Apache SparkUnderstand how Spark turns your code into jobs, stages and tasks, and why partitions, shuffles and data skew drive performance.
- KafkaKafka is a distributed log used for streaming data. Learn topics, partitions, consumer groups and delivery semantics before building streaming pipelines.
- AirflowAirflow schedules and orchestrates pipelines as DAGs. Learn scheduling, task dependencies, retries and idempotent task design.
- SnowflakeSnowflake is a cloud data warehouse that separates storage from compute. Learn virtual warehouses, micro-partitions, pruning, caching and cost control.