Compression inside Parquet

After encoding, each page is compressed on its own with the codec named in its column chunk's metadata, so a reader decompresses only the pages it needs. The format defines UNCOMPRESSED, SNAPPY, GZIP, LZO, BROTLI, ZSTD and LZ4_RAW (replacing the deprecated LZ4). This script rewrites the orders with each codec pyarrow 129 supports:

codecs.py: one file, six codecsPython
import os
import pyarrow.parquet as pq
from options import OPTIONS
table = pq.read_table("orders.parquet")
for codec in ("none", "snappy", "lz4", "gzip", "brotli", "zstd"):
    pq.write_table(table, f"codec-{codec}.parquet", compression=codec, **OPTIONS)
    print(f"{codec:7} {os.path.getsize(f'codec-{codec}.parquet'):>9,} bytes")
    os.remove(f"codec-{codec}.parquet")
Output
none      960,012 bytes
snappy    928,388 bytes
lz4       925,962 bytes
gzip      788,625 bytes
brotli    789,464 bytes
zstd      798,714 bytes

The uncompressed file is already 24 times smaller than the JSON Lines because the encodings did the heavy work; the best codec removes only another 18 percent. Snappy 6,616 and LZ4 barely help on well-encoded data, while gzip, Brotli and zstd 126 land within 1.3 percent of each other. Snappy is still pyarrow's and Spark 129 's default because it decompresses cheaply; Compression Codecs Compared measures the speed side that makes zstd the usual choice for new tables.