Apache Flink's Streaming Model

Apache Flink 129 (https://github.com/apache/flink 26,378 ) (Apache 2.0) is a distributed engine for stateful stream and batch processing. A JobManager turns each job into a dataflow graph and schedules its parallel operator instances into slots on TaskManagers. Flink tracks event time with watermarks, assertions that no records older than a timestamp are still expected; a window fires when the watermark passes its end, and watermarks merge by taking the minimum across input partitions. State lives in a state backend (heap or RocksDB, plus the disaggregated ForSt backend in 2.x) and is snapshotted by periodic checkpoints, which together with Kafka 129 transactions give exactly-once results. Flink 2.0 (March 2025) removed the DataSet API and Scala APIs; 2.3.0 (25 June 2026) adds SQL changelog conversion operators and more flexible materialized tables.

Flink runs in three ways. A session cluster, as in Flink 2.x Analytics, accepts many jobs; an application cluster runs one job's main() on its own JobManager; and the Flink Kubernetes 5,150 Operator manages either as custom resources, with savepoints for upgrades. The SQL client, the SQL Gateway (REST and JDBC) and the Table API in Java or Python submit SQL; the DataStream API in Java builds jobs by hand, much like Kafka Streams 129 ' DSL.

Why Section 6.10.7's daily windows dropped late lines and a Flink watermark would not
Why A BookNest Streams Application's daily windows dropped late lines and a Flink watermark would not