Hidden Partitioning

Hidden Partitioning and Partition Transforms

A Hive-style table partitions by a real column such as order_date that writers must fill and queries must filter on. Iceberg 129 partitions by a transform of a source column, recorded in the partition spec: identity, year, month, day, hour, bucket[N] (a hash into N buckets, for high-cardinality keys) and truncate[W] (string prefixes or rounded numbers). Queries filter on the source column, and the planner applies the transform itself. The script reads Iceberg's scan metrics for four queries on orders, partitioned by months(order_ts):

pruning.py: how many manifests and files four filters readPython
"""Hidden partitioning: filter on order_ts and Iceberg skips whole months by itself."""
from lake import spark
spark.conf.set("spark.sql.adaptive.enabled", "false")     # keeps the scan node at the leaf
def scan(label, where):                       # run a query, then read its scan metrics
    df = spark.sql(f"SELECT count(*) AS n, sum(total) FROM booknest.orders WHERE {where}")
    n = df.collect()[0].n
    m = df._jdf.queryExecution().executedPlan().collectLeaves().head().metrics()
    g = lambda k: m.get(k).get().value()
    print(f"{label:<15}{n:>7} rows  manifests {g('scannedDataManifests')} read "
          f"{g('skippedDataManifests')} skipped  files {g('resultDataFiles'):>2} read "
          f"{g('skippedDataFiles'):>2} skipped")
scan("all orders", "true")
scan("March 2026", "order_ts >= '2026-03-01' AND order_ts < '2026-04-01'")
scan("holidays", "order_ts BETWEEN '2025-12-24' AND '2026-01-06'")
scan("one customer", "customer_id = 4242")              # not a partition source column
Output
all orders      100000 rows  manifests 2 read 0 skipped  files 18 read  0 skipped
March 2026        5674 rows  manifests 1 read 1 skipped  files  1 read  5 skipped
holidays          2429 rows  manifests 2 read 0 skipped  files  2 read 16 skipped
one customer        12 rows  manifests 2 read 0 skipped  files 18 read  0 skipped

Pruning works at two levels. The manifest list's partition ranges ruled out the whole 2025 manifest for March 2026, and partition values and column bounds then skipped files within the manifests that were read. Random customer IDs leave column bounds too wide to prune anything. Choose a granularity that keeps files large (hundreds of megabytes in production); over-partitioning makes the small files of Compaction and Expiry.