A decompression bomb is a small input that expands until memory runs out. DEFLATE tops out near 1,000:1, so zip bombs nest archives; columnar files need no nesting, since run-length encoding stores millions of equal values in a few bytes:
import os, zlib
import pyarrow as pa, pyarrow.parquet as pq
bomb = zlib.compress(bytes(200_000_000), 9) # 200 MB of zero bytes
inflater = zlib.decompressobj()
chunk = inflater.decompress(bomb, 1_000_000) # max_length caps the output
print(f"zlib: {len(bomb):,} bytes; bounded read gave {len(chunk):,} bytes,",
f"{len(inflater.unconsumed_tail):,} input bytes left unread")
ones = pa.repeat(pa.scalar(1, pa.int64()), 50_000_000) # one value, 50 million times
pq.write_table(pa.table({"qty": ones}), "bomb.parquet", compression="zstd",
row_group_size=50_000_000, max_rows_per_page=50_000_000)
meta, size = pq.ParquetFile("bomb.parquet").metadata, os.path.getsize("bomb.parquet")
print(f"parquet: {size:,} bytes on disk, {meta.num_rows:,} rows, footer's",
f"uncompressed size {meta.row_group(0).total_byte_size:,} bytes")
table = pq.read_table("bomb.parquet")
print(f"decoded in memory: {table.nbytes:,} bytes ({table.nbytes // size:,}:1)")Output
zlib: 194,409 bytes; bounded read gave 1,000,000 bytes, 193,424 input bytes left unread parquet: 1,513 bytes on disk, 50,000,000 rows, footer's uncompressed size 994 bytes decoded in memory: 406,250,000 bytes (268,506:1)
The 1.5 KB Parquet 129 file decoded to 406 MB. Its footer's "uncompressed size" counts encoded bytes, so a check on it passes; check the row count instead, which the footer states exactly. For streams and ZIP entries, decompress in bounded steps, as max_length did, and abort past a limit.