RoadmapsPlan 6 of 9
Study plan · Plan 6 of 9
DSA for Data Engineers: Focused 100–200 Pattern Roadmap
A focused data structures and algorithms roadmap for Data Engineering interviews: the patterns that recur, in order, practised across roughly 100 to 200 problems.
Data Engineering interviews usually test DSA at an easy-to-medium level, with a bias toward problems that look like data processing. You do not need hundreds of hard puzzles. You need the recurring patterns, practised until they are automatic.
How to practise
- Learn the pattern, then solve problems in order of difficulty.
- State the brute force first, then optimise, then give time and space complexity.
- Write clean Python: clear names, small helper functions, edge cases (empty input, duplicates, ties).
- Re-solve problems you got wrong a week later.
The ranges per stage add up to roughly 100 to 200 problems in total. Stop a stage early once problems feel routine.
The plan
Stage 1: Arrays, strings and hashing
Frequency counts, grouping, deduplication and lookups with dicts and sets.
- Typical effort
- 20–30 problems
- Outcome
- You default to a hash map when you need lookups and can state the complexity.
Stage 2: Two pointers and sliding windows
Pairs in sorted data, longest substring or subarray with a constraint, moving aggregates.
- Typical effort
- 20–30 problems
- Outcome
- You recognise window problems and keep them O(n).
Stage 3: Sorting, intervals and merging
Merge overlapping intervals, meeting rooms, merging sorted streams.
- Typical effort
- 15–25 problems
- Outcome
- You can reason about interval overlap, which also appears in sessionisation and SCD logic.
Stage 4: Heaps and top-K
Top K frequent items, K-way merge, running medians.
- Typical effort
- 10–20 problems
- Outcome
- You use a heap for top-K and streaming problems.
Stage 5: Stacks and queues
Matching brackets, monotonic stacks, BFS with deques.
- Typical effort
- 10–15 problems
- Outcome
- You choose a deque for queues and know why list.pop(0) is slow.
Stage 6: Binary search
Search in sorted data and on the answer space.
- Typical effort
- 10–15 problems
- Outcome
- You write bug-free boundary conditions.
Stage 7: Graphs and trees (basics)
BFS, DFS, topological sort for dependencies (the same idea behind DAG schedulers).
- Typical effort
- 15–25 problems
- Outcome
- You can order tasks with dependencies and detect cycles.
Stage 8: Data-processing problems
Parse logs, aggregate records, deduplicate events, join two lists, in Python and SQL.
- Typical effort
- 15–30 problems
- Outcome
- You solve 'process these records' problems cleanly in both languages.