The two writers run here landed all 392,537 events exactly once, and Flink 129 's sink gives the same guarantee; they differ in what they cost and what they can do on the way in:
| Question | Kafka Connect 129 sink | Flink | Spark 129 Structured Streaming |
|---|---|---|---|
| Transformations | Single-message transforms only | Full SQL, joins, windows, state | Full SQL and Python, micro-batches |
| Upserts by key | Append only; MERGE downstream | Upsert mode, equality deletes | MERGE inside foreachBatch |
| Latency | Commit interval, minutes typical | Checkpoint interval, seconds possible | Trigger interval, seconds to minutes |
| What you operate | A Connect cluster | JobManager, TaskManagers, state | A Spark job or schedule |
| Exactly-once via | Kafka transactions, control topic | Checkpoints, two-phase commit | Checkpoint plus snapshot epoch |
For BookNest the Connect sink wins: the events need no transformation, a Connect cluster already runs Debezium 317,608 (CDC with Debezium 3.x), and one JSON file configures it. Spark Structured Streaming fits when landing also cleans or aggregates; Flink when latency must be seconds or the logic is stateful. Either way, plan compaction from day one and land raw events append-only, so a bug never rewrites the only copy of the history.