Study planner
90-Day Data Engineer Study Plan
90 days, 348 planned hours: about 3 hours on weekdays and 6 at weekends. Each day pairs one topic from every skill with a mock interview, revision or applications. Every topic links to its lesson.
Days completed0 of 90 days
Set a start date to see a calendar date on each day and highlight today. Saved in this browser only.
Phase 1: Foundations
Week 1SQL joins/windows solid + 15 DSA done
- Day 1SQL + DSA focus3 h
- SQLSELECT, WHERE & basic filtering
- DSAContains Duplicate
- PySparkRDD fundamentals
- SnowflakeThree-layer architecture
- KafkaTopic fundamentals
- AirflowDAG fundamentals
- AWSS3 buckets & objects
- System DesignDesign a Data Lake on cloud object storage
- AlsoRevise notes / flashcards
- Day 2Big Data Engineering3 h
- SQLDISTINCT and de-duplication
- DSAValid Anagram
- PySparkRDD transformations vs actions
- SnowflakeStorage layer (micro-partitions)
- KafkaPartitions & parallelism
- AirflowDAG authoring best practices
- AWSStorage classes
- System DesignDesign a Lakehouse (bronze/silver/gold)
- AlsoRevise notes / flashcards
- Day 3Cloud & Warehousing3 h
- SQLORDER BY multi-column sorting
- DSATwo Sum
- PySparkLazy evaluation & DAG
- SnowflakeCompute layer (virtual warehouses)
- KafkaPartition keys & ordering
- AirflowTasks & dependencies
- AWSLifecycle policies
- System DesignDesign a CDC pipeline from OLTP to warehouse
- AlsoApply to 3 target companies
- Day 4Streaming Systems3 h
- SQLLIMIT / TOP / FETCH FIRST
- DSAGroup Anagrams
- PySparkSparkSession & SparkContext
- SnowflakeCloud services layer
- KafkaReplication factor
- AirflowTaskFlow API
- AWSVersioning
- System DesignDesign a real-time streaming platform
- AlsoRevise notes / flashcards
- Day 5System Design Deep-Dive6 h
- SQLAggregate functions (SUM, AVG, MIN, MAX)
- DSATop K Frequent Elements
- PySparkSpark architecture (driver/executor)
- SnowflakeSeparation of storage & compute
- KafkaPartition assignment
- AirflowDynamic DAG generation
- AWSS3 partitioning for analytics
- System DesignDesign an event-driven architecture
- AlsoRevise notes / flashcards
- Day 6Data Modeling & SQL6 h
- SQLCOUNT vs COUNT(DISTINCT)
- DSAProduct of Array Except Self
- PySparkCluster managers (YARN/K8s/Standalone)
- SnowflakeMetadata management
- KafkaLog segments
- AirflowDynamic task mapping
- AWSS3 Select
- System DesignDesign an analytics / BI platform
- AlsoApply to 3 target companies
- Day 7Mock Interview Day3 h
- SQLGROUP BY fundamentals
- DSAValid Sudoku
- PySparkPartitions & parallelism
- SnowflakeMulti-cluster shared data
- KafkaRetention policies
- AirflowDAG params & templating (Jinja)
- AWSEvent notifications
- System DesignDesign a Customer 360 platform
- AlsoFull mock interview + review
Week 2Advanced SQL + 30 DSA + Spark core
- Day 8Big Data Engineering3 h
- SQLHAVING vs WHERE
- DSAEncode and Decode Strings
- PySparkNarrow vs wide transformations
- SnowflakeSnowflake editions
- KafkaCompaction (log cleanup)
- AirflowCatchup & start_date pitfalls
- AWSEncryption (SSE-S3/KMS)
- System DesignDesign a fraud detection pipeline
- AlsoRevise notes / flashcards
- Day 9Cloud & Warehousing3 h
- SQLAliases and column naming
- DSALongest Consecutive Sequence
- PySparkShuffle internals
- SnowflakePricing model (credits)
- KafkaTopic configuration
- AirflowBashOperator
- AWSAccess points
- System DesignDesign a recommendation data pipeline
- AlsoApply to 3 target companies
- Day 10Streaming Systems3 h
- SQLComparison & logical operators
- DSAMaximum Subarray
- PySparkmap vs flatMap
- SnowflakeCaching layers (result/local/metadata)
- KafkaProducer API
- AirflowPythonOperator
- AWSMultipart upload
- System DesignDesign a feature store
- AlsoRevise notes / flashcards
- Day 11System Design Deep-Dive3 h
- SQLIN, BETWEEN, LIKE patterns
- DSAMerge Intervals
- PySparkreduceByKey vs groupByKey
- SnowflakeVirtual warehouse sizing
- Kafkaacks (0/1/all)
- AirflowCustom operators
- AWSIAM users/groups/roles
- System DesignDesign a metrics / KPI platform
- AlsoRevise notes / flashcards
- Day 12Data Modeling & SQL6 h
- SQLNULL handling with IS NULL / COALESCE
- DSAInsert Interval
- PySparkmapPartitions
- SnowflakeScaling up vs scaling out
- KafkaIdempotent producer
- AirflowHooks
- AWSPolicies (identity/resource)
- System DesignDesign an ELT pipeline with dbt
- AlsoApply to 3 target companies
- Day 13SQL + DSA focus6 h
- SQLCASE WHEN expressions
- DSANon-overlapping Intervals
- PySparkPersistence & caching levels
- SnowflakeMulti-cluster warehouses
- KafkaBatching & linger.ms
- AirflowProvider packages
- AWSLeast privilege
- System DesignDesign a batch ingestion framework
- AlsoRevise notes / flashcards
- Day 14Mock Interview Day3 h
- SQLString functions (CONCAT, SUBSTRING, TRIM)
- DSARotate Image
- PySparkCheckpointing
- SnowflakeAuto-suspend & auto-resume
- KafkaCompression (snappy/lz4/zstd)
- AirflowKubernetesPodOperator
- AWSAssume role & STS
- System DesignDesign a data quality framework
- AlsoFull mock interview + review
Week 3PySpark performance + Snowflake basics
- Day 15Cloud & Warehousing3 h
- SQLDate functions basics
- DSASpiral Matrix
- PySparkAccumulators
- SnowflakeWarehouse concurrency
- KafkaPartitioner strategies
- AirflowBranching (BranchPythonOperator)
- AWSCross-account access
- System DesignDesign a data observability system
- AlsoApply to 3 target companies
- Day 16Streaming Systems3 h
- SQLCAST and data type conversion
- DSASet Matrix Zeroes
- PySparkBroadcast variables
- SnowflakeQuery queuing
- KafkaRetries & delivery
- AirflowTrigger rules
- AWSService-linked roles
- System DesignDesign a data catalog & lineage system
- AlsoRevise notes / flashcards
- Day 17System Design Deep-Dive3 h
- SQLArithmetic & rounding
- DSASubarray Sum Equals K
- PySparkRepartition vs coalesce
- SnowflakeWorkload isolation
- Kafkamax.in.flight & ordering
- AirflowSensor basics
- AWSPermission boundaries
- System DesignDesign a GDPR/PII compliant pipeline
- AlsoRevise notes / flashcards
- Day 18Data Modeling & SQL3 h
- SQLUNION vs UNION ALL
- DSAMajority Element
- PySparkSpark UI & stages/tasks
- SnowflakeResource monitors
- KafkaProducer buffering
- AirflowPoke vs reschedule mode
- AWSPolicy evaluation logic
- System DesignDesign a clickstream analytics pipeline
- AlsoApply to 3 target companies
- Day 19SQL + DSA focus6 h
- SQLINNER JOIN basics
- DSAValid Palindrome
- PySparkJob/stage/task hierarchy
- SnowflakeMicro-partitions deep dive
- KafkaConsumer API
- AirflowExternalTaskSensor
- AWSLambda fundamentals
- System DesignDesign an IoT sensor data pipeline
- AlsoRevise notes / flashcards
- Day 20Big Data Engineering6 h
- SQLLEFT JOIN basics
- DSATwo Sum II Input Array Is Sorted
- PySparkClosures & serialization
- SnowflakeNatural clustering
- KafkaPoll loop
- AirflowFileSensor
- AWSTriggers & event sources
- System DesignDesign a log ingestion & search platform
- AlsoRevise notes / flashcards
- Day 21Mock Interview Day3 h
- SQLRIGHT JOIN & FULL OUTER JOIN
- DSA3Sum
- PySparkDataFrame API basics
- SnowflakeClustering keys
- KafkaOffset management
- AirflowDeferrable operators & triggers
- AWSConcurrency & scaling
- System DesignDesign a near-real-time dashboard backend
- AlsoFull mock interview + review
Week 41 end-to-end project + 50 DSA
- Day 22Streaming Systems3 h
- SQLSelf Join patterns
- DSAContainer With Most Water
- PySparkDataset vs DataFrame
- SnowflakeClustering depth & information
- KafkaAuto vs manual commit
- AirflowSmart sensors
- AWSCold starts
- System DesignDesign a slowly changing dimension framework
- AlsoRevise notes / flashcards
- Day 23System Design Deep-Dive3 h
- SQLCross Join & cartesian products
- DSATrapping Rain Water
- PySparkSchema definition (StructType)
- SnowflakeReclustering
- KafkaConsumer lag
- AirflowScheduler internals
- AWSLayers
- System DesignDesign an idempotent reprocessing system
- AlsoRevise notes / flashcards
- Day 24Data Modeling & SQL3 h
- SQLMulti-table joins
- DSAMove Zeroes
- PySparkselect / withColumn / filter
- SnowflakeWhen to cluster
- KafkaSeek & rewind
- AirflowSchedule intervals & cron
- AWSLambda + S3/Kinesis
- System DesignDesign a multi-tenant data platform
- AlsoApply to 3 target companies
- Day 25SQL + DSA focus3 h
- SQLAnti-join (NOT EXISTS / LEFT JOIN NULL)
- DSASort Colors
- PySparkgroupBy & aggregations
- SnowflakeSearch optimization service
- KafkaDeserialization
- AirflowTimetables
- AWSError handling & DLQ
- System DesignDesign a cost-optimized warehouse strategy
- AlsoRevise notes / flashcards
- Day 26Big Data Engineering6 h
- SQLSemi-join with EXISTS
- DSABest Time to Buy and Sell Stock
- PySparkJoins in Spark SQL
- SnowflakeStream object basics
- Kafkafetch.min/max.bytes
- AirflowData-aware scheduling (Datasets)
- AWSMemory/timeout tuning
- System DesignDesign a data mesh architecture
- AlsoRevise notes / flashcards
- Day 27Cloud & Warehousing6 h
- SQLSubqueries in WHERE
- DSALongest Substring Without Repeating Characters
- PySparkWindow functions in Spark
- SnowflakeStandard vs append-only streams
- KafkaConsumer group basics
- AirflowExecution date vs logical date
- AWSGlue Data Catalog
- System DesignDesign a streaming ETL with Kafka + Spark
- AlsoApply to 3 target companies
- Day 28Mock Interview Day3 h
- SQLSubqueries in FROM (derived tables)
- DSALongest Repeating Character Replacement
- PySparkUDFs and pandas UDFs
- SnowflakeInsert-only streams
- KafkaGroup coordinator
- AirflowSLAs
- AWSGlue Crawlers
- System DesignDesign a near-zero downtime migration
- AlsoFull mock interview + review
Week 5Kafka + streaming fundamentals
- Day 29System Design Deep-Dive3 h
- SQLScalar subqueries
- DSAPermutation in String
- PySparkCatalyst optimizer
- SnowflakeStream offsets
- KafkaRebalancing protocols
- AirflowXCom push/pull
- AWSGlue ETL jobs (Spark)
- System DesignDesign an A/B testing data pipeline
- AlsoRevise notes / flashcards
- Day 30Data Modeling & SQL3 h
- SQLCorrelated subqueries
- DSAMinimum Window Substring
- PySparkTungsten execution engine
- SnowflakeCDC with streams
- KafkaCooperative sticky assignor
- AirflowCustom XCom backends
- AWSGlue Studio
- System DesignDesign a marketing attribution pipeline
- AlsoApply to 3 target companies
Month 1: Core SQL+DSA+Spark mastered; 1 project; start applying Target: 60% prep • first OAs & recruiter calls.
Phase 2: Advanced + Projects
Week 5Kafka + streaming fundamentals
- Day 31SQL + DSA focus3 h
- SQLCommon Table Expressions (CTEs)
- DSASliding Window Maximum
- PySparkReading/writing Parquet
- SnowflakeStale streams
- KafkaStatic membership
- AirflowVariables & Connections
- AWSDynamic frames
- System DesignDesign a financial reconciliation pipeline
- AlsoRevise notes / flashcards
- Day 32Big Data Engineering3 h
- SQLMultiple chained CTEs
- DSAMaximum Average Subarray I
- PySparkReading/writing JSON & CSV
- SnowflakeTask fundamentals
- KafkaPartition reassignment
- AirflowParams
- AWSGlue bookmarks
- System DesignDesign a time-series metrics store
- AlsoRevise notes / flashcards
- Day 33Cloud & Warehousing6 h
- SQLConditional aggregation
- DSAValid Parentheses
- PySparkHandling nested/complex types
- SnowflakeScheduled tasks (CRON)
- KafkaLeader & followers
- AirflowTask state lifecycle
- AWSGlue triggers/workflows
- System DesignDesign a search indexing pipeline
- AlsoApply to 3 target companies
- Day 34Streaming Systems6 h
- SQLPIVOT rows to columns
- DSAMin Stack
- PySparkexplode & posexplode
- SnowflakeTask trees / DAGs
- KafkaIn-Sync Replicas (ISR)
- AirflowSequentialExecutor
- AWSGlue partitions
- System DesignDesign a notification/alerting pipeline
- AlsoRevise notes / flashcards
- Day 35Mock Interview Day3 h
- SQLUNPIVOT columns to rows
- DSAEvaluate Reverse Polish Notation
- PySparkpivot in Spark
- SnowflakeServerless tasks
- Kafkamin.insync.replicas
- AirflowLocalExecutor
- AWSPySpark on Glue
- System DesignDesign an order-events processing system
- AlsoFull mock interview + review
Week 6Airflow + AWS Glue/Redshift
- Day 36Data Modeling & SQL3 h
- SQLGROUPING SETS
- DSAGenerate Parentheses
- PySparkTemp views & global views
- SnowflakeTask + Stream pattern
- KafkaUnclean leader election
- AirflowCeleryExecutor
- AWSGlue cost tuning
- System DesignDesign a payment events pipeline (exactly-once)
- AlsoApply to 3 target companies
- Day 37SQL + DSA focus3 h
- SQLROLLUP and CUBE
- DSADaily Temperatures
- PySparkSQL vs DataFrame performance
- SnowflakeTask error handling
- KafkaHigh watermark
- AirflowKubernetesExecutor
- AWSAthena basics
- System DesignDesign a ride-hailing surge pricing pipeline
- AlsoRevise notes / flashcards
- Day 38Big Data Engineering3 h
- SQLWindow function basics (OVER)
- DSACar Fleet
- PySparkColumn expressions
- SnowflakeSnowpipe continuous ingestion
- KafkaReplica fetcher
- AirflowCeleryKubernetes hybrid
- AWSQuerying S3 with SQL
- System DesignDesign a social media feed analytics system
- AlsoRevise notes / flashcards
- Day 39Cloud & Warehousing3 h
- SQLPARTITION BY clause
- DSALargest Rectangle in Histogram
- PySparkHandling nulls (na.fill/drop)
- SnowflakeAuto-ingest with notifications
- KafkaRack awareness
- AirflowParallelism & pools
- AWSPartition projection
- System DesignDesign a video streaming analytics pipeline
- AlsoApply to 3 target companies
- Day 40Streaming Systems6 h
- SQLROW_NUMBER()
- DSANext Greater Element I
- PySparkType casting & schema evolution
- SnowflakeSnowpipe REST API
- KafkaKRaft vs ZooKeeper
- AirflowBackfills & reruns
- AWSCTAS & external tables
- System DesignDesign an e-commerce inventory sync system
- AlsoRevise notes / flashcards
- Day 41System Design Deep-Dive6 h
- SQLRANK() vs DENSE_RANK()
- DSAImplement Stack using Queues
- PySparkUnderstanding shuffle cost
- SnowflakeCOPY INTO command
- KafkaAt-most-once
- AirflowClearing tasks
- AWSAthena performance tuning
- System DesignDesign a unified batch + streaming (Lambda/Kappa)
- AlsoRevise notes / flashcards
- Day 42Mock Interview Day3 h
- SQLNTILE() bucketing
- DSAImplement Queue using Stacks
- PySparkSpark partitioning strategy
- SnowflakeFile formats & staging
- KafkaAt-least-once
- AirflowRetries & alerting
- AWSWorkgroups
- System DesignDesign a data warehouse for a SaaS product
- AlsoFull mock interview + review
Week 7Data modeling + 80 DSA
- Day 43SQL + DSA focus3 h
- SQLLAG() and LEAD()
- DSADesign Circular Queue
- PySparkAQE (Adaptive Query Execution)
- SnowflakeInternal vs external stages
- KafkaExactly-once semantics (EOS)
- AirflowLogging configuration
- AWSFederated queries
- System DesignDesign a churn prediction data pipeline
- AlsoRevise notes / flashcards
- Day 44Big Data Engineering3 h
- SQLFIRST_VALUE / LAST_VALUE
- DSANumber of Recent Calls
- PySparkDynamic partition pruning
- SnowflakeSnowpipe Streaming
- KafkaTransactions in Kafka
- AirflowMetrics (StatsD)
- AWSCost per query
- System DesignDesign a real-time leaderboard
- AlsoRevise notes / flashcards
- Day 45Cloud & Warehousing3 h
- SQLRunning totals with SUM OVER
- DSABinary Search
- PySparkPredicate & projection pushdown
- SnowflakeTime Travel basics
- KafkaIdempotence end-to-end
- AirflowAirflow REST API
- AWSRedshift architecture
- System DesignDesign a geospatial analytics pipeline
- AlsoApply to 3 target companies
- Day 46Streaming Systems3 h
- SQLMoving average with frame clause
- DSASearch a 2D Matrix
- PySparkCaching strategy
- SnowflakeAT / BEFORE clauses
- KafkaRead-process-write pattern
- AirflowConnections & secrets backend
- AWSDistribution styles
- System DesignDesign an ad-bidding analytics pipeline
- AlsoRevise notes / flashcards
- Day 47System Design Deep-Dive6 h
- SQLPercent of total
- DSAKoko Eating Bananas
- PySparkBucketing in Spark
- SnowflakeUNDROP
- KafkaConnect framework
- AirflowIdempotent tasks
- AWSSort keys
- System DesignDesign a backfill & late-data handling system
- AlsoRevise notes / flashcards
- Day 48Data Modeling & SQL6 h
- SQLCumulative distribution
- DSAFind Minimum in Rotated Sorted Array
- PySparkSalting for skew
- SnowflakeRetention period config
- KafkaSource vs sink connectors
- AirflowAvoiding top-level code
- AWSVacuum & analyze
- System DesignDesign a schema registry & contract system
- AlsoApply to 3 target companies
- Day 49Mock Interview Day3 h
- SQLDate truncation & bucketing
- DSASearch in Rotated Sorted Array
- PySparkDetecting data skew
- SnowflakeFail-safe (7 days)
- KafkaStandalone vs distributed
- AirflowTesting DAGs
- AWSRedshift Spectrum
- System DesignDesign a self-serve analytics platform
- AlsoFull mock interview + review
Week 82nd project + system design start
- Day 50Big Data Engineering3 h
- SQLGenerate date series / calendar table
- DSATime Based Key-Value Store
- PySparkSpill to disk diagnosis
- SnowflakeCloning with Time Travel
- KafkaConverters & SMTs
- AirflowCI/CD for DAGs
- AWSConcurrency scaling
- System DesignDesign an LLM/RAG data ingestion pipeline
- AlsoRevise notes / flashcards
- Day 51Cloud & Warehousing3 h
- SQLRecursive CTEs
- DSAMedian of Two Sorted Arrays
- PySparkMemory management (unified)
- SnowflakeRBAC model
- KafkaOffset storage
- AirflowResource management
- AWSRA3 & managed storage
- System DesignDesign a vector embeddings pipeline
- AlsoApply to 3 target companies
- Day 52Streaming Systems3 h
- SQLHierarchical / org-chart queries
- DSAFind First and Last Position
- PySparkGC tuning
- SnowflakeRoles & role hierarchy
- KafkaDead letter queue
- AirflowScaling Airflow
- AWSWorkload management (WLM)
- System DesignDesign a data SLA & freshness monitoring system
- AlsoRevise notes / flashcards
- Day 53System Design Deep-Dive3 h
- SQLGraph traversal in SQL
- DSAReverse Linked List
- PySparkspark.sql.shuffle.partitions tuning
- SnowflakeObject ownership
- KafkaDebezium CDC
- AirflowDAG fundamentals
- AWSMaterialized views
- System DesignDesign a Data Lake on cloud object storage
- AlsoRevise notes / flashcards
- Day 54Data Modeling & SQL6 h
- SQLGaps and Islands problem
- DSAMerge Two Sorted Lists
- PySparkExecutor sizing
- SnowflakeColumn-level security / masking
- KafkaStreams DSL
- AirflowDAG authoring best practices
- AWSCOPY & UNLOAD
- System DesignDesign a Lakehouse (bronze/silver/gold)
- AlsoApply to 3 target companies
- Day 55SQL + DSA focus6 h
- SQLSessionization of events
- DSALinked List Cycle
- PySparkCores & memory config
- SnowflakeRow access policies
- KafkaKStream vs KTable
- AirflowTasks & dependencies
- AWSEMR clusters
- System DesignDesign a CDC pipeline from OLTP to warehouse
- AlsoRevise notes / flashcards
- Day 56Mock Interview Day3 h
- SQLRetention analysis
- DSAReorder List
- PySparkBroadcast hash join
- SnowflakeNetwork policies
- KafkaStateful processing
- AirflowTaskFlow API
- AWSSpark on EMR
- System DesignDesign a real-time streaming platform
- AlsoFull mock interview + review
Week 95 system design scenarios + mocks
- Day 57Cloud & Warehousing3 h
- SQLCohort analysis
- DSARemove Nth Node From End of List
- PySparkSort-merge join tuning
- SnowflakeData encryption
- KafkaWindowing
- AirflowDynamic DAG generation
- AWSEMR on EKS
- System DesignDesign an event-driven architecture
- AlsoApply to 3 target companies
- Day 58Streaming Systems3 h
- SQLFunnel / conversion analysis
- DSACopy List with Random Pointer
- PySparkJoin strategy selection
- SnowflakeOAuth & SSO
- KafkaJoins in Streams
- AirflowDynamic task mapping
- AWSInstance fleets & spot
- System DesignDesign an analytics / BI platform
- AlsoRevise notes / flashcards
- Day 59System Design Deep-Dive3 h
- SQLMedian & percentile (PERCENTILE_CONT)
- DSAAdd Two Numbers
- PySparkAvoiding wide transformations
- SnowflakeSecure views
- KafkaState stores
- AirflowDAG params & templating (Jinja)
- AWSBootstrap actions
- System DesignDesign a Customer 360 platform
- AlsoRevise notes / flashcards
- Day 60Data Modeling & SQL3 h
- SQLTop-N per group
- DSAFind the Duplicate Number
- PySparkRepartition before write
- SnowflakeCredit consumption analysis
- KafkaInteractive queries
- AirflowCatchup & start_date pitfalls
- AWSEMR Serverless
- System DesignDesign a fraud detection pipeline
- AlsoApply to 3 target companies
Month 2: Cloud+Streaming+Modeling done; 2 projects; active interviews Target: 85% prep • multiple onsites in pipeline.
Phase 3: Interview Mastery
Week 95 system design scenarios + mocks
- Day 61SQL + DSA focus6 h
- SQLDeduplicate keeping latest record
- DSALRU Cache
- PySparkFile size optimization
- SnowflakeWarehouse right-sizing
- KafkaProcessor API
- AirflowBashOperator
- AWSCost optimization on EMR
- System DesignDesign a recommendation data pipeline
- AlsoRevise notes / flashcards
- Day 62Big Data Engineering6 h
- SQLYear-over-year growth
- DSAMerge k Sorted Lists
- PySparkSmall files problem
- SnowflakeAuto-suspend tuning
- KafkaExactly-once in Streams
- AirflowPythonOperator
- AWSKinesis Data Streams
- System DesignDesign a feature store
- AlsoRevise notes / flashcards
- Day 63Mock Interview Day3 h
- SQLMonth-over-month change
- DSAReverse Nodes in k-Group
- PySparkColumn pruning
- SnowflakeQuery result caching
- KafkaSchema Registry basics
- AirflowCustom operators
- AWSShards & throughput
- System DesignDesign a metrics / KPI platform
- AlsoFull mock interview + review
Week 10110 DSA + revise weak areas
- Day 64Streaming Systems3 h
- SQLRolling 7/30 day metrics
- DSAInvert Binary Tree
- PySparkCaching vs recompute trade-off
- SnowflakeAvoiding spilling
- KafkaAvro/Protobuf/JSON schemas
- AirflowHooks
- AWSKinesis Firehose
- System DesignDesign an ELT pipeline with dbt
- AlsoRevise notes / flashcards
- Day 65System Design Deep-Dive3 h
- SQLFirst & last touch attribution
- DSAMaximum Depth of Binary Tree
- PySparkCoalesce on output
- SnowflakeStorage cost (time travel/fail-safe)
- KafkaSchema evolution & compatibility
- AirflowProvider packages
- AWSKinesis vs Kafka
- System DesignDesign a batch ingestion framework
- AlsoRevise notes / flashcards
- Day 66Data Modeling & SQL3 h
- SQLMarket basket / co-occurrence
- DSADiameter of Binary Tree
- PySparkSkew join optimization
- SnowflakeResource monitors for cost
- KafkaSubject naming strategies
- AirflowKubernetesPodOperator
- AWSEnhanced fan-out
- System DesignDesign a data quality framework
- AlsoApply to 3 target companies
- Day 67SQL + DSA focus3 h
- SQLSlowly changing dimension queries
- DSABalanced Binary Tree
- PySparkCost-based optimization
- SnowflakeMaterialized view costs
- KafkaSerializers/deserializers
- AirflowBranching (BranchPythonOperator)
- AWSKCL/KPL
- System DesignDesign a data observability system
- AlsoRevise notes / flashcards
- Day 68Big Data Engineering6 h
- SQLDetecting consecutive streaks
- DSASame Tree
- PySparkStructured Streaming model
- SnowflakeSecure Data Sharing
- KafkaSchema references
- AirflowTrigger rules
- AWSKinesis Data Analytics
- System DesignDesign a data catalog & lineage system
- AlsoRevise notes / flashcards
- Day 69Cloud & Warehousing6 h
- SQLPivoting dynamic columns
- DSASubtree of Another Tree
- PySparkInput sources (Kafka/files)
- SnowflakeReader accounts
- KafkaCluster sizing
- AirflowSensor basics
- AWSState machines
- System DesignDesign a GDPR/PII compliant pipeline
- AlsoApply to 3 target companies
- Day 70Mock Interview Day3 h
- SQLConditional window frames
- DSABinary Tree Level Order Traversal
- PySparkOutput sinks & modes
- SnowflakeData Marketplace
- KafkaMonitoring & JMX metrics
- AirflowPoke vs reschedule mode
- AWSStandard vs Express
- System DesignDesign a clickstream analytics pipeline
- AlsoFull mock interview + review
Week 11Full mock loops + behavioral prep
- Day 71System Design Deep-Dive3 h
- SQLEXCEPT / INTERSECT set operations
- DSABinary Tree Right Side View
- PySparkTriggers & micro-batch
- SnowflakeListings & exchanges
- KafkaThroughput tuning
- AirflowExternalTaskSensor
- AWSMap & Parallel states
- System DesignDesign an IoT sensor data pipeline
- AlsoRevise notes / flashcards
- Day 72Data Modeling & SQL3 h
- SQLLATERAL / CROSS APPLY joins
- DSACount Good Nodes in Binary Tree
- PySparkWatermarking
- SnowflakeCross-region/cloud sharing
- KafkaQuotas
- AirflowFileSensor
- AWSError handling & retries
- System DesignDesign a log ingestion & search platform
- AlsoApply to 3 target companies
- Day 73SQL + DSA focus3 h
- SQLJSON parsing in SQL
- DSAConstruct Tree from Preorder and Inorder
- PySparkEvent-time vs processing-time
- SnowflakeShares & grants
- KafkaSecurity (SASL/SSL/ACL)
- AirflowDeferrable operators & triggers
- AWSIntegration with Glue/Lambda
- System DesignDesign a near-real-time dashboard backend
- AlsoRevise notes / flashcards
- Day 74Big Data Engineering3 h
- SQLArray & nested data handling
- DSABinary Tree Maximum Path Sum
- PySparkStateful aggregations
- SnowflakeZero-copy cloning
- KafkaMirror Maker 2
- AirflowSmart sensors
- AWSOrchestration patterns
- System DesignDesign a slowly changing dimension framework
- AlsoRevise notes / flashcards
- Day 75Cloud & Warehousing6 h
- SQLRegex matching in SQL
- DSASerialize and Deserialize Binary Tree
- PySparkStream-stream joins
- SnowflakeTransient & temporary tables
- KafkaDisaster recovery
- AirflowScheduler internals
- AWSMetrics & alarms
- System DesignDesign an idempotent reprocessing system
- AlsoApply to 3 target companies
- Day 76Streaming Systems6 h
- SQLQuery optimization & cost reduction
- DSALowest Common Ancestor of Binary Tree
- PySparkStream-static joins
- SnowflakeExternal tables
- KafkaTiered storage
- AirflowSchedule intervals & cron
- AWSLogs & Log Insights
- System DesignDesign a multi-tenant data platform
- AlsoRevise notes / flashcards
- Day 77Mock Interview Day3 h
- SQLReading EXPLAIN / execution plans
- DSAValidate Binary Search Tree
- PySparkCheckpointing in streaming
- SnowflakeSemi-structured (VARIANT)
- KafkaCruise Control rebalancing
- AirflowTimetables
- AWSEventBridge rules
- System DesignDesign a cost-optimized warehouse strategy
- AlsoFull mock interview + review
Week 12Final revision + apply aggressively
- Day 78Data Modeling & SQL3 h
- SQLIndex design (B-tree, composite)
- DSAKth Smallest Element in a BST
- PySparkExactly-once in streaming
- SnowflakeFLATTEN function
- KafkaTombstones & log compaction
- AirflowData-aware scheduling (Datasets)
- AWSDashboards
- System DesignDesign a data mesh architecture
- AlsoApply to 3 target companies
- Day 79SQL + DSA focus3 h
- SQLCovering indexes
- DSALowest Common Ancestor of a BST
- PySparkHandling late data
- SnowflakeDynamic tables
- KafkaProducer/consumer interceptors
- AirflowExecution date vs logical date
- AWSMonitoring pipelines
- System DesignDesign a streaming ETL with Kafka + Spark
- AlsoRevise notes / flashcards
- Day 80Big Data Engineering3 h
- SQLIndex selectivity & cardinality
- DSAInsert into a BST
- PySparkForeach & foreachBatch
- SnowflakeIceberg tables
- KafkaTopic fundamentals
- AirflowSLAs
- AWSCustom metrics
- System DesignDesign a near-zero downtime migration
- AlsoRevise notes / flashcards
- Day 81Cloud & Warehousing3 h
- SQLPartitioning strategies
- DSADelete Node in a BST
- PySparkBackpressure & rate limiting
- SnowflakeThree-layer architecture
- KafkaPartitions & parallelism
- AirflowXCom push/pull
- AWSLake Formation basics
- System DesignDesign an A/B testing data pipeline
- AlsoApply to 3 target companies
- Day 82Streaming Systems6 h
- SQLPartition pruning
- DSAConvert Sorted Array to BST
- PySparkDelta Lake architecture
- SnowflakeStorage layer (micro-partitions)
- KafkaPartition keys & ordering
- AirflowCustom XCom backends
- AWSFine-grained access control
- System DesignDesign a marketing attribution pipeline
- AlsoRevise notes / flashcards
- Day 83System Design Deep-Dive6 h
- SQLMaterialized views
- DSAImplement Trie (Prefix Tree)
- PySparkACID transactions on Delta
- SnowflakeCompute layer (virtual warehouses)
- KafkaReplication factor
- AirflowVariables & Connections
- AWSLF-Tags
- System DesignDesign a financial reconciliation pipeline
- AlsoRevise notes / flashcards
- Day 84Mock Interview Day3 h
- SQLQuery rewriting for performance
- DSADesign Add and Search Words
- PySparkTransaction log (_delta_log)
- SnowflakeCloud services layer
- KafkaPartition assignment
- AirflowParams
- AWSGoverned tables
- System DesignDesign a time-series metrics store
- AlsoFull mock interview + review
Week 13Offer negotiation & decision
- Day 85SQL + DSA focus3 h
- SQLAvoiding full table scans
- DSAWord Search II
- PySparkTime travel & versioning
- SnowflakeSeparation of storage & compute
- KafkaLog segments
- AirflowTask state lifecycle
- AWSBlueprints
- System DesignDesign a search indexing pipeline
- AlsoRevise notes / flashcards
- Day 86Big Data Engineering3 h
- SQLJoin algorithms (hash, merge, nested loop)
- DSAKth Largest Element in a Stream
- PySparkMERGE / upsert
- SnowflakeMetadata management
- KafkaRetention policies
- AirflowSequentialExecutor
- AWSCross-account data sharing
- System DesignDesign a notification/alerting pipeline
- AlsoRevise notes / flashcards
- Day 87Cloud & Warehousing3 h
- SQLStatistics & the optimizer
- DSALast Stone Weight
- PySparkSchema enforcement
- SnowflakeMulti-cluster shared data
- KafkaCompaction (log cleanup)
- AirflowLocalExecutor
- AWSVPC & networking basics
- System DesignDesign an order-events processing system
- AlsoApply to 3 target companies
- Day 88Streaming Systems3 h
- SQLDeadlocks & locking
- DSAK Closest Points to Origin
- PySparkSchema evolution
- SnowflakeSnowflake editions
- KafkaTopic configuration
- AirflowCeleryExecutor
- AWSDynamoDB for DE
- System DesignDesign a payment events pipeline (exactly-once)
- AlsoRevise notes / flashcards
- Day 89System Design Deep-Dive6 h
- SQLIsolation levels (ACID)
- DSAKth Largest Element in an Array
- PySparkOPTIMIZE & compaction
- SnowflakePricing model (credits)
- KafkaProducer API
- AirflowKubernetesExecutor
- AWSRDS/Aurora basics
- System DesignDesign a ride-hailing surge pricing pipeline
- AlsoRevise notes / flashcards
- Day 90Data Modeling & SQL6 h
- SQLMVCC concepts
- DSATask Scheduler
- PySparkZ-Ordering
- SnowflakeCaching layers (result/local/metadata)
- Kafkaacks (0/1/all)
- AirflowCeleryKubernetes hybrid
- AWSSQS & SNS
- System DesignDesign a social media feed analytics system
- AlsoApply to 3 target companies
Month 3: Interview-ready across stack; multiple offers; negotiate offers Target: 100% prep • offers in hand.