Partition Evolution

Evolving BookNest's Partition Spec Without Rewriting Data

Suppose June's sales grow enough to deserve daily partitions. A Hive 129 table would need a full copy; in Iceberg 129 it is a metadata change: a new partition spec gets the next ID, new writes use it, and old files keep the spec they were written with. Here the spec moves to days(order_ts), and only June is rewritten into it:

evolve_spec.py: months to days, with only June rewrittenPython
"""Partition evolution: months -> days; old files keep the spec they were written with."""
from lake import spark
spark.sql("""ALTER TABLE booknest.orders
             REPLACE PARTITION FIELD order_ts_month WITH days(order_ts) AS order_day""")
spark.sql("""CALL lake.system.rewrite_data_files(table => 'booknest.orders',   -- June only
             where => "order_ts >= '2026-06-01'", options => map('rewrite-all', 'true'))""")
for r in spark.sql("""SELECT spec_id, count(*) AS files, sum(record_count) AS orders,
        min(partition.order_ts_month) AS month, min(partition.order_day) AS day
        FROM booknest.orders.files GROUP BY spec_id ORDER BY spec_id""").collect():
    print(f"spec {r.spec_id}: {r.files} files, {r.orders} orders, first partition "
          f"{r.month if r.day is None else r.day}")
Output
spec 0: 17 files, 94502 orders, first partition 660
spec 1: 30 files, 5498 orders, first partition 2026-06-01

Seventeen monthly files from January 2025 (month 660) to May 2026 kept spec 0, June became 30 daily files under spec 1, and the orders still add up to 100,000. The planner prunes each file with its own spec, so queries do not change. Without the rewrite_data_files call, only new appends would be daily, the usual way to evolve.