Batch Processing with Apache Spark processed BookNest's order history in nightly batches. Some questions cannot wait for the night: is this payment fraudulent, which book is about to sell out? Answering them means handling each order event as it happens and handing it to every system that needs it, without wiring each producer to each consumer.
Apache Kafka 129 is the standard tool for that job: a replicated, append-only log that many applications read at their own pace, plus Connect, schema registries, Kafka Streams 129 and Flink 129 around it. Since 4.0 it needs no ZooKeeper.
You run Kafka 4.3 in Docker 514 , first as one broker and then as a six-node cluster, and push BookNest's order events through every layer. The second half compares managed Kafka services with their real prices, and runs BookNest on Aiven's free plan.
What you will learn
How Kafka's partitioned log works on disk, and how replication and the KRaft quorum keep it available.
How to run Kafka 4.x in Docker and manage topics, producers and consumer groups.
How to write reliable producers and consumers, use transactions, and choose consumer or share groups.
How Connect, Schema Registry, Kafka Streams and Flink build pipelines on the log.
How to monitor, secure and tune a cluster, and how to choose and migrate between managed services.
Sections
- Event Streaming Concepts
- The Kafka Log Under the Hood
- Kafka 4.x in Docker
- Topics and CLI Tools
- Producers and Delivery
- Transactions and Exactly-Once
- Consumers and Queues
- Kafka Connect
- Schema Registry
- Kafka Streams
- ksqlDB and Apache Flink
- Kafka Clients
- Monitoring and Kafka UIs
- Securing a Kafka Cluster
- Tuning and Tiered Storage
- Confluent Cloud in Depth
- MSK, Event Hubs, Google Kafka
- Aiven, Redpanda and Others
- Choosing and Migrating
- Kafka Alternatives
- Test Yourself!