After encoding, each page is compressed on its own with the codec named in its column chunk's metadata, so a reader decompresses only the pages it needs. The format defines UNCOMPRESSED, SNAPPY, GZIP, LZO, BROTLI, ZSTD and LZ4_RAW (replacing the deprecated LZ4). This script rewrites the orders with each codec pyarrow 129 supports:
import os
import pyarrow.parquet as pq
from options import OPTIONS
table = pq.read_table("orders.parquet")
for codec in ("none", "snappy", "lz4", "gzip", "brotli", "zstd"):
pq.write_table(table, f"codec-{codec}.parquet", compression=codec, **OPTIONS)
print(f"{codec:7} {os.path.getsize(f'codec-{codec}.parquet'):>9,} bytes")
os.remove(f"codec-{codec}.parquet")none 960,012 bytes snappy 928,388 bytes lz4 925,962 bytes gzip 788,625 bytes brotli 789,464 bytes zstd 798,714 bytes
The uncompressed file is already 24 times smaller than the JSON Lines because the encodings did the heavy work; the best codec removes only another 18 percent. Snappy 6,616 and LZ4 barely help on well-encoded data, while gzip, Brotli and zstd 126 land within 1.3 percent of each other. Snappy is still pyarrow's and Spark 129 's default because it decompresses cheaply; Compression Codecs Compared measures the speed side that makes zstd the usual choice for new tables.