A lakehouse bill has three meters (us-east-1 list prices, 2 October 2026): storage, $0.023 per GB-month on S3 Standard; requests, $0.005 per 1,000 PUTs or LISTs and $0.0004 per 1,000 GETs; and compute, such as $0.44 per DPU-hour (4 vCPUs, 16 GB) for AWS Glue 24 Spark 129 jobs, plus $0.09 per GB of egress. Small files hit requests and compute: costs/requests.sh counts MinIO 30,943 's requests (mc admin trace) during a DuckDB 61,228 scan, here over orders written as 400 files and again after compaction:
# The same DuckDB scan over 400 small files, then over the compacted table.
S="$DEMOS/scripts"
bash "$S/spark.sh" costs small_files.py make
bash "$DEMOS/costs/requests.sh" orders_small
bash "$S/spark.sh" costs small_files.py compact
bash "$DEMOS/costs/requests.sh" orders_small
bash "$S/spark.sh" costs small_files.py dropbooknest.orders_small data files: 400 orders_small: 400 requests on data files, 3 on metadata, 1698 ms rewrite_data_files: 400 files -> 1 booknest.orders_small data files: 1 orders_small: 3 requests on data files, 4 on metadata, 406 ms
The same 100,000 rows cost 133 times as many requests in 400 files and took 4.2 times as long (2.4 times in an earlier run; this 4-CPU host is shared, so compare ratios, not milliseconds). For a dashboard that runs this scan 10,000 times a day, that is 120 million GETs a month, $48, against 0.9 million, $0.36. The bigger loss is compute: each file adds a round trip, a footer to parse and a task to schedule, billed by the second. Compaction (Compaction and Expiry) costs one job run.