Menu

Advanced project · Project 7 of 8

Real-Time Analytics Pipeline

Build a pipeline that turns a stream of order events into per-minute revenue and order counts by category, visible on a dashboard within a minute, and correct even when events arrive late.

  • Advanced
  • Kafka · Spark Structured Streaming · A serving store (PostgreSQL works for a project) · A simple dashboard (notebook or lightweight web chart)
  • 2 min read
  • Updated Oct 2026

Requirements

  • Produce synthetic order events with event timestamps, including some late events
  • Aggregate into one-minute event-time windows with a watermark
  • Upsert results into a serving table
  • Show a live chart of the last hour
  • Recover after restarts without double counting

Technology stack

Kafka, Spark Structured Streaming, A serving store (PostgreSQL works for a project), A simple dashboard (notebook or lightweight web chart)

Dataset

Synthetic events from a producer script that deliberately sends some events late and some twice.

Business context

Operations teams want to see within a minute when orders spike or stop. The interesting engineering is not the chart; it is producing correct numbers when events are late or duplicated.

Architecture

  1. Producer sends order events with event time and event id.
  2. Spark reads from Kafka, deduplicates within the watermark.
  3. One-minute event-time windows by category.
  4. Upserts into the serving table.
  5. Dashboard polls the serving table.
Late events update the right window until the watermark passes.

Compare your streaming numbers with a batch recount over the same raw events. Any difference must be explained (for example, events dropped after the watermark).

By Data Career Hub Editorial · Last reviewed Oct 2026

Progress is saved in this browser only. No account needed.

Search
Filter by type