Menu

Course · Cloud

AWS

AWS for Data Engineers: S3, IAM, Lambda, Glue, Athena, Redshift, EMR, Kinesis, Step Functions, Lake Formation and how they fit into one data platform.

Lessons
8
Interview questions
0
Projects & case studies
2
Reading time
~2 h

About this course

Amazon Web Services is the most common cloud in Data Engineer job descriptions. An AWS data platform is usually a lake in Amazon S3, described by the AWS Glue Data Catalog, loaded by streaming and batch ingestion, transformed with Spark on Glue or EMR, queried with Athena or Redshift, orchestrated with Step Functions or Airflow, governed with IAM and Lake Formation, and watched with CloudWatch.

Start with the map of the AWS data stack, then S3 and IAM, because every other service depends on them. After that, follow the lessons in order or jump to the service your team uses.

Your progress

Saved in this browser only

Practise

Course structure

Lessons

Work through the lessons in order. Completed lessons show a tick; lessons you have opened are outlined.

Start here

The complete overview of the course in one read.

  1. AWS for Data Engineers: The AWS Data Stack and How the Services Fit TogetherA map of the AWS data stack for Data Engineers: ingest, store, process, query, orchestrate, govern and monitor, with a reference architecture and how to choose.Beginner12 min

Beginner

Core concepts you will use every day.

  1. Amazon S3 for Data EngineersS3 for data lakes: buckets and keys, storage classes, lifecycle rules, versioning, partitioned layouts, events, encryption, access points and multipart upload.Beginner23 min
  2. AWS IAM for Data Engineers: Roles, Policies and Cross-Account AccessIAM for data pipelines: users, groups and roles, identity and resource policies, least privilege, STS, cross-account access, boundaries and policy evaluation.Beginner24 min

Intermediate

Patterns used in production pipelines.

  1. AWS Lambda for Data PipelinesUse AWS Lambda in data pipelines: triggers, S3 and Kinesis processing, concurrency, cold starts, layers, memory and timeout tuning, and error handling with DLQs.Intermediate17 min
  2. AWS Glue: Data Catalog, Crawlers and Spark ETL JobsAWS Glue for Data Engineers: the Data Catalog, crawlers, partitions, Spark ETL jobs, DynamicFrames, Glue Studio, job bookmarks, triggers, workflows and cost tuning.Intermediate18 min
  3. Amazon Athena: Serverless SQL on S3Query S3 with Amazon Athena: external tables, CTAS, partition projection, performance tuning, workgroups, federated queries and what drives the cost per query.Intermediate17 min
  4. Amazon Redshift for Data EngineersRedshift for Data Engineers: architecture, RA3, distribution and sort keys, COPY and UNLOAD, vacuum, WLM, concurrency scaling, Spectrum and Serverless.Intermediate21 min

Advanced

Performance, internals and edge cases.

  1. Amazon EMR: Spark Clusters, EMR on EKS and EMR ServerlessRun Spark on Amazon EMR: cluster node types, Spark tuning, instance fleets and Spot, bootstrap actions, EMR on EKS, EMR Serverless and how to keep EMR costs down.Advanced15 min

Projects and case studies

Apply what you learned and prepare material to discuss in interviews.

Projects

System design case studies

Resources

Related courses

Plan your learning

Search
Filter by type