Iceberg, Delta and Hudi

Iceberg, Delta Lake and Hudi Side by Side

All three store data in Parquet 129 on object storage and commit by atomically adding metadata (What the Formats Share); they differ in how that metadata is organized and what it optimizes for. The table summarizes what this chapter ran:

The three open table formats compared
Aspect Apache Iceberg 129 Delta Lake 237,929 Apache Hudi 129
Origin Netflix, 2017 Databricks 2,717 , 2019 Uber, 2016
Metadata Metadata JSON, manifest tree JSON log plus checkpoints Timeline plus metadata table
Commit point Catalog pointer swap Next log file, put-if-absent Timeline instant, with locks
Row changes Position, equality deletes, DVs Deletion vectors, rewrite COW rewrite or MOR log files
Partitioning Hidden transforms, evolvable Columns, liquid clustering Columns, expression indexes
Key lookups Predicate scan Predicate scan Record-level and other indexes
Engines Spark 129 , Trino 403,499 , Flink 129 , DuckDB 61,228 , Snowflake Spark, Databricks, Fabric, delta-rs Spark, Flink, Trino (read)
Catalogs REST spec, many servers Unity Catalog, metastores Metastores
Version run here 1.12.0 4.4.0, delta-rs 1.6.6 1.2.1

The differences in practice: Iceberg's hidden partitioning and REST catalogs make it the easiest format to share across engines; Delta's single log is simple and fastest inside Databricks; Hudi's indexes and table services suit upsert-heavy streams. All three are Apache-2.0, and all three now read deletion vectors or log-style deltas, so their capabilities converge while their ecosystems differ.