Retention and GDPR Deletion

Retention Policies and GDPR Deletion in an Immutable Format

A retention policy states how long each kind of data may be kept: orders for the years tax law demands, tickets for two years, clickstream for 90 days. GDPR Article 17 adds the right to erasure (Article 12 allows a month to act), the CCPA a right to delete, and Malaysia's PDPA a retention principle: data no longer needed must be destroyed. In a table format a DELETE never edits a file. With merge-on-read, BookNest's choice for orders, the current snapshot still points at the old file plus a delete file, and older snapshots keep it for time travel. privacy/erase.py erases customer 12 from a copy of orders, counting rows in the table, by time travel and in the Parquet 129 files themselves (read by DuckDB 61,228 in privacy/erase_check.py):

privacy/erase.py: erase one customer and check the bytes, not just the tablePython
"""Erase customer 12 from an Iceberg table, then check the bytes in MinIO, not the table."""
from functools import partial
from lake import spark
from erase_check import report        # counts rows in the table, by time travel, in raw files
T = "booknest.orders_gdpr"
spark.sql(f"""CREATE OR REPLACE TABLE {T} USING iceberg
              TBLPROPERTIES ('write.delete.mode' = 'merge-on-read')
              AS SELECT * FROM booknest.orders""")      # a scratch copy of BookNest's orders
snap = spark.sql(f"SELECT snapshot_id FROM {T}.snapshots").first()[0]
files = [r[0] for r in
         spark.sql(f"SELECT DISTINCT _file FROM {T} WHERE customer_id = 12").collect()]
check = partial(report, spark, T, 12, snap, files)
check("before")
spark.sql(f"DELETE FROM {T} WHERE customer_id = 12")
check("DELETE")
r = spark.sql(f"""CALL lake.system.rewrite_data_files(table => '{T}',
                  options => map('delete-file-threshold', '1'))""").first()
print(f"rewrote {r.rewritten_data_files_count} data file(s) into {r.added_data_files_count}")
check("rewrite_data_files")
spark.sql(f"""CALL lake.system.expire_snapshots(table => '{T}',
              older_than => current_timestamp(), retain_last => 1)""")
check("expire_snapshots")
spark.sql(f"DROP TABLE {T} PURGE")
Output
before              table 15 | time travel 15 | raw Parquet 15
DELETE              table 0 | time travel 15 | raw Parquet 15
rewrote 12 data file(s) into 1
rewrite_data_files  table 0 | time travel 15 | raw Parquet 15
expire_snapshots    table 0 | time travel Cannot find snapshot with ID | raw Parquet deleted

After the DELETE the customer vanished from every query, yet all 15 orders sat in MinIO 30,943 . Rewriting the 12 files with deletes cleaned the current snapshot; only expire_snapshots dropped the old snapshots and their files. Run both in the erasure job, then chase other copies: tags and branches (Branches and Tags) pin snapshots, MinIO versioning (Consistency and Versioning) keeps old versions, and backups and Kafka 129 topics hold their own.

Retention is cheaper: a partition-aligned DELETE drops whole files in a metadata-only commit, and history.expire.max-snapshot-age-ms bounds what expire_snapshots keeps. Crypto-shredding, encrypting each customer's fields with their own key and deleting the key, helps where physical deletion everywhere is impractical.