A streaming job runs continuously. It reads events from a durable log such as an Apache Kafka 129 topic, updates its results as each event or small group of events arrives, and never finishes. Fraud checks on a payment, low-stock alerts, live sales counters and recommendations that react to the last few clicks all need it.
Streaming buys freshness, measured in seconds, at the price of new concepts and harder operations:
Event time versus processing time. An event carries the time it happened, but arrives later, sometimes much later and out of order. Results grouped by event time must decide how long to wait for stragglers; a watermark is the job's estimate that "no more events older than this are coming".
Windows and state. "Orders per genre in the last hour" needs a window and must remember counts between events. That state has to survive a crash, so engines checkpoint it.
Delivery guarantees. After a failure, events may be processed again. Most pipelines get at-least-once delivery and make their writes idempotent, for example by ignoring an order_id already stored. Exactly-once results are possible within Kafka and engines such as Flink 129 , but end to end they depend on every external system in the pipeline (Transactions and Exactly-Once).
Always on. A stream job is a service: it needs monitoring, capacity for peaks and a plan for upgrades.
Engines differ in how they process the stream. Kafka Streams 129 and Apache Flink handle one event at a time; Spark 129 Structured Streaming runs a rapid series of small batches, called micro-batches. Apache Kafka and Managed Cloud Kafka covers Kafka, Kafka Streams and Flink, and Structured Streaming looks at Spark's streaming API.