A Hive-style table partitions by a real column such as order_date that writers must fill and queries must filter on. Iceberg 129 partitions by a transform of a source column, recorded in the partition spec: identity, year, month, day, hour, bucket[N] (a hash into N buckets, for high-cardinality keys) and truncate[W] (string prefixes or rounded numbers). Queries filter on the source column, and the planner applies the transform itself. The script reads Iceberg's scan metrics for four queries on orders, partitioned by months(order_ts):
"""Hidden partitioning: filter on order_ts and Iceberg skips whole months by itself."""
from lake import spark
spark.conf.set("spark.sql.adaptive.enabled", "false") # keeps the scan node at the leaf
def scan(label, where): # run a query, then read its scan metrics
df = spark.sql(f"SELECT count(*) AS n, sum(total) FROM booknest.orders WHERE {where}")
n = df.collect()[0].n
m = df._jdf.queryExecution().executedPlan().collectLeaves().head().metrics()
g = lambda k: m.get(k).get().value()
print(f"{label:<15}{n:>7} rows manifests {g('scannedDataManifests')} read "
f"{g('skippedDataManifests')} skipped files {g('resultDataFiles'):>2} read "
f"{g('skippedDataFiles'):>2} skipped")
scan("all orders", "true")
scan("March 2026", "order_ts >= '2026-03-01' AND order_ts < '2026-04-01'")
scan("holidays", "order_ts BETWEEN '2025-12-24' AND '2026-01-06'")
scan("one customer", "customer_id = 4242") # not a partition source columnall orders 100000 rows manifests 2 read 0 skipped files 18 read 0 skipped March 2026 5674 rows manifests 1 read 1 skipped files 1 read 5 skipped holidays 2429 rows manifests 2 read 0 skipped files 2 read 16 skipped one customer 12 rows manifests 2 read 0 skipped files 18 read 0 skipped
Pruning works at two levels. The manifest list's partition ranges ruled out the whole 2025 manifest for March 2026, and partition values and column bounds then skipped files within the manifests that were read. Random customer IDs leave column bounds too wide to prune anything. Choose a granularity that keeps files large (hundreds of megabytes in production); over-partitioning makes the small files of Compaction and Expiry.