In practice the surprises come from what nobody deletes and what nobody turns off. costs/stored.sh compares, for each table, the data bytes its current snapshot needs (Trino 403,499 's $files table) with everything stored under its prefix in MinIO 30,943 (mc du):
# Data bytes each table's current snapshot needs versus all bytes under its prefix in MinIO.
printf "%-16s %10s %10s %8s\n" table live_MiB stored_MiB objects
for t in customers daily_sales order_events order_items orders; do
live=$(docker exec l1-trino trino --execute 2>/dev/null \
"SELECT sum(file_size_in_bytes) FROM iceberg.booknest.\"$t\$files\"" | tr -d '"')
read -r bytes objs <<< "$(mc du --json lake/warehouse/booknest/$t |
python3 -c 'import json,sys; d=json.load(sys.stdin); print(d["size"], d["objects"])')"
printf "%-16s %10.2f %10.2f %8s\n" $t "$(bc -l <<< "$live/1048576")" \
"$(bc -l <<< "$bytes/1048576")" $objs
donetable live_MiB stored_MiB objects customers 0.12 0.50 16 daily_sales 0.05 0.99 264 order_events 3.36 6.89 102 order_items 0.49 0.63 16 orders 1.03 1.62 85
Across these tables only 5.1 of 10.6 MiB is live data. daily_sales stores 20 times its size: every CREATE OR REPLACE and MERGE of Query Engines on the Lakehouse kept its snapshot, files and metadata. customers holds four copies, because the REST fixture's DROP TABLE left the files of earlier loads behind, an orphaned copy of personal data too. At 50 TB the same ratio is a five-figure annual line. The usual culprits: unexpired snapshots and orphan files (schedule Iceberg in Production's maintenance), clusters left running overnight, small files, egress and cross-zone traffic, and tables copied elsewhere "temporarily". Tag every bucket and cluster with an owner and a cost center.