Kafka 129 earns its operational cost when many independent consumers read and replay the same high-volume history and feed it to Connect, Debezium 317,608 , Flink 129 and the lakehouse (Lakehouses, Data Quality and Governance), as with BookNest's orders. Many needs differ:
| If you mainly need | Consider | Why |
|---|---|---|
| A few thousand background jobs a day | A database queue, SQS | Nothing new to run |
| Request-reply between services | NATS | Replies and subjects built in |
| Edge sites and devices | NATS leaf nodes | 20 MiB, one binary |
| Per-message routing, priorities, TTL | RabbitMQ 28,807 queues | Exchanges and policies |
| Millions of topics, many tenants | Pulsar | Tenants, cheap topics |
| Serverless, inside one cloud | SQS, Pub/Sub, Service Bus | No cluster at all |
A PostgreSQL 1,289 table read with SELECT ... FOR UPDATE SKIP LOCKED is a good queue at modest volume and commits each job with the business data; share groups (Consumers and Queues) made Kafka a queue, but not a cheap one. Cloud queues trade replay for zero operations: Amazon SQS keeps a message 4 days by default and 14 at most, up to 1 MiB, and cannot replay a deleted one; Google Cloud Pub/Sub 1 replays only if you enable retention and seek.
The warnings run both ways: a small system that grows into the event backbone means a migration later (A Migration Runbook), and a second messaging system doubles the monitoring and security work of Monitoring and Kafka UIs-Securing a Kafka Cluster. BookNest keeps Kafka, and would add NATS only if its warehouse devices needed a broker on site.