Streaming Processing

Streaming Processing and Its Trade-offs

A streaming job runs continuously. It reads events from a durable log such as an Apache Kafka 129 topic, updates its results as each event or small group of events arrives, and never finishes. Fraud checks on a payment, low-stock alerts, live sales counters and recommendations that react to the last few clicks all need it.

Streaming buys freshness, measured in seconds, at the price of new concepts and harder operations:

Engines differ in how they process the stream. Kafka Streams 129 and Apache Flink handle one event at a time; Spark 129 Structured Streaming runs a rapid series of small batches, called micro-batches. Apache Kafka and Managed Cloud Kafka covers Kafka, Kafka Streams and Flink, and Structured Streaming looks at Spark's streaming API.