A retention policy states how long each kind of data may be kept: orders for the years tax law demands, tickets for two years, clickstream for 90 days. GDPR Article 17 adds the right to erasure (Article 12 allows a month to act), the CCPA a right to delete, and Malaysia's PDPA a retention principle: data no longer needed must be destroyed. In a table format a DELETE never edits a file. With merge-on-read, BookNest's choice for orders, the current snapshot still points at the old file plus a delete file, and older snapshots keep it for time travel. privacy/erase.py erases customer 12 from a copy of orders, counting rows in the table, by time travel and in the Parquet 129 files themselves (read by DuckDB 61,228 in privacy/erase_check.py):
"""Erase customer 12 from an Iceberg table, then check the bytes in MinIO, not the table."""
from functools import partial
from lake import spark
from erase_check import report # counts rows in the table, by time travel, in raw files
T = "booknest.orders_gdpr"
spark.sql(f"""CREATE OR REPLACE TABLE {T} USING iceberg
TBLPROPERTIES ('write.delete.mode' = 'merge-on-read')
AS SELECT * FROM booknest.orders""") # a scratch copy of BookNest's orders
snap = spark.sql(f"SELECT snapshot_id FROM {T}.snapshots").first()[0]
files = [r[0] for r in
spark.sql(f"SELECT DISTINCT _file FROM {T} WHERE customer_id = 12").collect()]
check = partial(report, spark, T, 12, snap, files)
check("before")
spark.sql(f"DELETE FROM {T} WHERE customer_id = 12")
check("DELETE")
r = spark.sql(f"""CALL lake.system.rewrite_data_files(table => '{T}',
options => map('delete-file-threshold', '1'))""").first()
print(f"rewrote {r.rewritten_data_files_count} data file(s) into {r.added_data_files_count}")
check("rewrite_data_files")
spark.sql(f"""CALL lake.system.expire_snapshots(table => '{T}',
older_than => current_timestamp(), retain_last => 1)""")
check("expire_snapshots")
spark.sql(f"DROP TABLE {T} PURGE")before table 15 | time travel 15 | raw Parquet 15 DELETE table 0 | time travel 15 | raw Parquet 15 rewrote 12 data file(s) into 1 rewrite_data_files table 0 | time travel 15 | raw Parquet 15 expire_snapshots table 0 | time travel Cannot find snapshot with ID | raw Parquet deleted
After the DELETE the customer vanished from every query, yet all 15 orders sat in MinIO 30,943 . Rewriting the 12 files with deletes cleaned the current snapshot; only expire_snapshots dropped the old snapshots and their files. Run both in the erasure job, then chase other copies: tags and branches (Branches and Tags) pin snapshots, MinIO versioning (Consistency and Versioning) keeps old versions, and backups and Kafka 129 topics hold their own.
Retention is cheaper: a partition-aligned DELETE drops whole files in a metadata-only commit, and history.expire.max-snapshot-age-ms bounds what expire_snapshots keeps. Crypto-shredding, encrypting each customer's fields with their own key and deleting the key, helps where physical deletion everywhere is impractical.