Study planner
PySpark tracker
Core Spark, Spark SQL, performance, optimisation, Structured Streaming and Delta Lake. 106 topics. Open a topic to study it, then tick your progress.
PySpark studied0 of 106
Studied, Practised, Revised, Mock and a confidence score (1–5) for each topic, as in the planner spreadsheet. A topic counts as complete once it is marked studied.
Saved in this browser only. Back it up from the study planner page.
| # | Topic | Category | Level | Studied | Practised | Revised | Mock | Confidence |
|---|---|---|---|---|---|---|---|---|
| 1 | Spark architecture (driver/executor) | Core Spark | Beginner | |||||
| 2 | Cluster managers (YARN/K8s/Standalone) | Core Spark | Intermediate | |||||
| 3 | SparkSession & SparkContext | Core Spark | Beginner | |||||
| 4 | RDD fundamentals | Core Spark | Beginner | |||||
| 5 | RDD transformations vs actions | Core Spark | Beginner | |||||
| 6 | map vs flatMap | Core Spark | Intermediate | |||||
| 7 | reduceByKey vs groupByKey | Core Spark | Advanced | |||||
| 8 | mapPartitions | Core Spark | Advanced | |||||
| 9 | Closures & serialization | Core Spark | Expert | |||||
| 10 | Lazy evaluation & DAG | Core Spark | Beginner | |||||
| 11 | Narrow vs wide transformations | Core Spark | Intermediate | |||||
| 12 | Job/stage/task hierarchy | Core Spark | Expert | |||||
| 13 | Spark UI & stages/tasks | Core Spark | Expert | |||||
| 14 | Partitions & parallelism | Core Spark | Intermediate | |||||
| 15 | Shuffle internals | Core Spark | Intermediate | |||||
| 16 | Understanding shuffle cost | Performance | Beginner | |||||
| 17 | Repartition vs coalesce | Core Spark | Expert | |||||
| 18 | Spark partitioning strategy | Performance | Beginner | |||||
| 19 | spark.sql.shuffle.partitions tuning | Performance | Expert | |||||
| 20 | Detecting data skew | Performance | Advanced | |||||
| 21 | Salting for skew | Performance | Advanced | |||||
| 22 | Skew join optimization | Optimization | Expert | |||||
| 23 | Persistence & caching levels | Core Spark | Advanced | |||||
| 24 | Caching strategy | Performance | Intermediate | |||||
| 25 | Caching vs recompute trade-off | Optimization | Advanced | |||||
| 26 | Checkpointing | Core Spark | Advanced | |||||
| 27 | Broadcast variables | Core Spark | Expert | |||||
| 28 | Accumulators | Core Spark | Advanced | |||||
| 29 | Catalyst optimizer | Spark SQL | Intermediate | |||||
| 30 | Tungsten execution engine | Spark SQL | Intermediate | |||||
| 31 | Whole-stage code generation | Deployment & Tuning | Advanced | |||||
| 32 | Cost-based optimization | Optimization | Expert | |||||
| 33 | Predicate & projection pushdown | Performance | Intermediate | |||||
| 34 | Column pruning | Optimization | Advanced | |||||
| 35 | SQL vs DataFrame performance | Spark SQL | Expert | |||||
| 36 | AQE (Adaptive Query Execution) | Performance | Beginner | |||||
| 37 | Dynamic partition pruning | Performance | Intermediate | |||||
| 38 | Memory management (unified) | Performance | Expert | |||||
| 39 | Executor sizing | Performance | Expert | |||||
| 40 | Cores & memory config | Performance | Expert | |||||
| 41 | GC tuning | Performance | Expert | |||||
| 42 | Spill to disk diagnosis | Performance | Expert | |||||
| 43 | Dynamic allocation | Deployment & Tuning | Beginner | |||||
| 44 | Speculative execution | Deployment & Tuning | Beginner | |||||
| 45 | Spark on Kubernetes | Deployment & Tuning | Beginner | |||||
| 46 | Monitoring with Spark History Server | Deployment & Tuning | Expert | |||||
| 47 | DataFrame API basics | Spark SQL | Beginner | |||||
| 48 | Dataset vs DataFrame | Spark SQL | Beginner | |||||
| 49 | Schema definition (StructType) | Spark SQL | Beginner | |||||
| 50 | select / withColumn / filter | Spark SQL | Beginner | |||||
| 51 | Column expressions | Spark SQL | Expert | |||||
| 52 | Handling nulls (na.fill/drop) | Spark SQL | Expert | |||||
| 53 | Type casting & schema evolution | Spark SQL | Expert | |||||
| 54 | groupBy & aggregations | Spark SQL | Beginner | |||||
| 55 | pivot in Spark | Spark SQL | Advanced | |||||
| 56 | Temp views & global views | Spark SQL | Expert | |||||
| 57 | Joins in Spark SQL | Spark SQL | Intermediate | |||||
| 58 | Broadcast hash join | Optimization | Beginner | |||||
| 59 | Sort-merge join tuning | Optimization | Beginner | |||||
| 60 | Join strategy selection | Optimization | Beginner | |||||
| 61 | Bucketing in Spark | Performance | Advanced | |||||
| 62 | Bucketed joins | Deployment & Tuning | Expert | |||||
| 63 | Window functions in Spark | Spark SQL | Intermediate | |||||
| 64 | UDFs and pandas UDFs | Spark SQL | Intermediate | |||||
| 65 | Pandas API on Spark | Deployment & Tuning | Intermediate | |||||
| 66 | Handling nested/complex types | Spark SQL | Advanced | |||||
| 67 | explode & posexplode | Spark SQL | Advanced | |||||
| 68 | Reading/writing Parquet | Spark SQL | Advanced | |||||
| 69 | Reading/writing JSON & CSV | Spark SQL | Advanced | |||||
| 70 | ORC & Avro formats | Deployment & Tuning | Advanced | |||||
| 71 | Handling corrupt records (badRecordsPath) | Deployment & Tuning | Intermediate | |||||
| 72 | Reading from JDBC sources | Deployment & Tuning | Intermediate | |||||
| 73 | Partition discovery | Deployment & Tuning | Advanced | |||||
| 74 | Repartition before write | Optimization | Intermediate | |||||
| 75 | Coalesce on output | Optimization | Expert | |||||
| 76 | File size optimization | Optimization | Intermediate | |||||
| 77 | Small files problem | Optimization | Advanced | |||||
| 78 | Avoiding wide transformations | Optimization | Intermediate | |||||
| 79 | Custom partitioners | Deployment & Tuning | Expert | |||||
| 80 | Structured Streaming model | Streaming | Beginner | |||||
| 81 | Input sources (Kafka/files) | Streaming | Beginner | |||||
| 82 | Output sinks & modes | Streaming | Beginner | |||||
| 83 | Triggers & micro-batch | Streaming | Intermediate | |||||
| 84 | Checkpointing in streaming | Streaming | Expert | |||||
| 85 | Foreach & foreachBatch | Streaming | Expert | |||||
| 86 | Backpressure & rate limiting | Streaming | Expert | |||||
| 87 | Event-time vs processing-time | Streaming | Intermediate | |||||
| 88 | Watermarking | Streaming | Intermediate | |||||
| 89 | Handling late data | Streaming | Expert | |||||
| 90 | Stateful aggregations | Streaming | Advanced | |||||
| 91 | Stream-stream joins | Streaming | Advanced | |||||
| 92 | Stream-static joins | Streaming | Advanced | |||||
| 93 | Exactly-once in streaming | Streaming | Expert | |||||
| 94 | Delta Lake architecture | Delta Lake | Beginner | |||||
| 95 | ACID transactions on Delta | Delta Lake | Beginner | |||||
| 96 | Transaction log (_delta_log) | Delta Lake | Beginner | |||||
| 97 | Time travel & versioning | Delta Lake | Intermediate | |||||
| 98 | Schema enforcement | Delta Lake | Intermediate | |||||
| 99 | Schema evolution | Delta Lake | Advanced | |||||
| 100 | MERGE / upsert | Delta Lake | Intermediate | |||||
| 101 | Change Data Feed | Delta Lake | Expert | |||||
| 102 | Streaming with Delta | Delta Lake | Expert | |||||
| 103 | OPTIMIZE & compaction | Delta Lake | Advanced | |||||
| 104 | Z-Ordering | Delta Lake | Advanced | |||||
| 105 | VACUUM & retention | Delta Lake | Expert | |||||
| 106 | Liquid clustering | Delta Lake | Expert |
No topics match these filters
Clear a filter to see more topics.