Iceberg 129 tracks columns by field ID, not by name or position, so a schema change only writes new metadata. The script adds a column, renames another, widens a third, then tries two unsafe changes:
"""Schema evolution is a metadata change: no data file is rewritten, no snapshot is added."""
from lake import spark
T = "booknest.orders"
state = lambda: spark.sql(f"""SELECT (SELECT count(*) FROM {T}.files) AS files,
(SELECT snapshot_id FROM {T}.refs WHERE name = 'main') AS main""").first()
spark.sql(f"ALTER TABLE {T} ADD COLUMN currency STRING COMMENT 'ISO 4217 code' AFTER total")
spark.sql(f"ALTER TABLE {T} RENAME COLUMN coupon TO coupon_code")
spark.sql(f"ALTER TABLE {T} ALTER COLUMN customer_id TYPE BIGINT") # int -> long: safe
print("after three changes:", state())
for change in ("total TYPE DECIMAL(12,3)", "order_id TYPE INT"): # unsafe: refused
try:
spark.sql(f"ALTER TABLE {T} ALTER COLUMN {change}")
except Exception as e: # Iceberg's or Spark's own check
j = getattr(e, "java_exception", None)
msg = (j.getMessage() if j else str(e)).replace("Unsupported table change: ", "")
print(f"{change}: {msg[:67]}")
spark.sql(f"ALTER TABLE {T} RENAME COLUMN coupon_code TO coupon") # back to the names
spark.sql(f"ALTER TABLE {T} DROP COLUMN currency") # later sections useOutput
after three changes: Row(files=47, main=4637335077536227122) total TYPE DECIMAL(12,3): Cannot change column type: total: decimal(10, 2) -> decimal(12, 3) order_id TYPE INT: [NOT_SUPPORTED_CHANGE_COLUMN] ALTER TABLE ALTER/CHANGE COLUMN is no
The table still has 47 files and the same main snapshot as in Branches and Tags: nothing was rewritten, old rows read currency as NULL, and coupon_code finds data written as coupon because the ID did not change. Iceberg refused the scale change and Spark 129 the narrowing, since either could corrupt values.
| Change | Allowed | Note |
|---|---|---|
| Add, drop, rename, reorder a column | Yes | Metadata only; IDs never reused |
| int to long, float to double | Yes | Widening only |
| Increase decimal precision | Yes | Scale must stay the same |
| Narrow a type, change scale or kind | No | Add a new column and backfill |
Version 3 tables add column defaults for old rows. The script restores the original names for later sections; customer_id stays a BIGINT.