Picking a Table Format

Picking a Table Format for BookNest

BookNest's data is read by Spark 129 , Trino 403,499 and DuckDB 61,228 (Query Engines on the Lakehouse), lands from Kafka 129 through Kafka Connect 129 (Streaming into the Lakehouse), and sits behind a REST catalog: Iceberg 129 is its primary format, and every later section uses it. Its orders matched across three engines, and its partitions evolved in place.

Delta Lake 237,929 still has a place: delta-rs wrote and merged the catalog table from plain Python in under a second, which suits small single-node jobs, and a Databricks 2,717 or Fabric consumer would read Delta natively. Rather than keep two copies, write such tables with UniForm or translate them with XTable, so Iceberg engines read the same files. Hudi 129 would earn a place only if BookNest's order updates grew into a high-volume, key-based CDC stream where record-level indexes and merge-on-read pay off.

The rule generalizes: pick the format your engines and catalog support best, keep one copy of the data, and use metadata translation for the exceptions. Changing formats later is a metadata migration (Iceberg's migrate and snapshot procedures, Delta's CONVERT TO DELTA, XTable), not a rewrite of every file, provided the data files are Parquet 129 .