Kafka courseLesson 3 of 6
Kafka course · Lesson 3 of 6
Kafka vs Queue-Based Messaging for Data Pipelines
When to use a distributed log like Kafka and when a traditional message queue fits better: retention, replay, ordering, fan-out, per-message acknowledgement and operations.
Both move messages between systems. They are built around different models, and that decides which fits a data pipeline.
| Distributed log (Kafka) | Message queue (typical) | |
|---|---|---|
| Model | Append-only log; consumers track their own position | Messages delivered to a consumer and removed after acknowledgement |
| Retention | Time- or size-based, independent of consumption | Usually until consumed (or expired) |
| Replay | Yes: reset the offset and read again | Generally no, once acknowledged |
| Many independent consumers | Natural: each consumer group reads everything | Needs fan-out (topics/exchanges) configured per subscriber |
| Ordering | Per partition | Varies; often per queue, weakened by multiple consumers |
| Work distribution | By partition | Per message, flexible |
| Typical strengths | High-throughput event streams, CDC, replayable pipelines | Task queues, per-message work, request decoupling |
Choose a log when
- Several systems need the same events (lakehouse, search, alerts).
- You need to replay history after a bug or to bootstrap a new consumer.
- Throughput is high and events are part of an ongoing stream.
- Per-key ordering matters (all changes for one order in sequence).
Choose a queue when
- Each message is a unit of work that one worker should do once (resize an image, send an email).
- You want per-message acknowledgement, retries and dead-lettering without managing offsets.
- Volumes are moderate and replay is not needed.
In data engineering
Most pipeline use cases (event streaming, CDC, feeding a lakehouse) fit the log model because replay and multiple consumers are valuable. Task orchestration between pipeline steps is usually handled by an orchestrator, not by either.
Common mistakes
- Using a queue where consumers later need to replay history.
- Using Kafka for a small task queue and taking on its operational overhead.
- Expecting global ordering from either.
Key takeaway
Logs retain and replay events for many consumers; queues distribute units of work and forget them. Pick by replay, fan-out and ordering needs.
Progress is saved in this browser only. No account needed.