Connect Sink or Flink?

Choosing Between the Kafka Connect Sink and Flink

The two writers run here landed all 392,537 events exactly once, and Flink 129 's sink gives the same guarantee; they differ in what they cost and what they can do on the way in:

Three ways to land a Kafka 129 topic in Iceberg 129
Question Kafka Connect 129 sink Flink Spark 129 Structured Streaming
Transformations Single-message transforms only Full SQL, joins, windows, state Full SQL and Python, micro-batches
Upserts by key Append only; MERGE downstream Upsert mode, equality deletes MERGE inside foreachBatch
Latency Commit interval, minutes typical Checkpoint interval, seconds possible Trigger interval, seconds to minutes
What you operate A Connect cluster JobManager, TaskManagers, state A Spark job or schedule
Exactly-once via Kafka transactions, control topic Checkpoints, two-phase commit Checkpoint plus snapshot epoch

For BookNest the Connect sink wins: the events need no transformation, a Connect cluster already runs Debezium 317,608 (CDC with Debezium 3.x), and one JSON file configures it. Spark Structured Streaming fits when landing also cleans or aggregates; Flink when latency must be seconds or the logic is stateful. Either way, plan compaction from day one and land raw events append-only, so a bug never rewrites the only copy of the history.