Vendor benchmarks use their own data; this one uses two of this chapter's files. orders.jsonl is 23.6 MB of row-oriented text, and orders.arrow (Arrow IPC and Feather Files) is 11.8 MB of uncompressed, decoded columns. Each codec compresses and decompresses each file three times, keeps the fastest run, checks the round trip, and reports speeds relative to gzip, since absolute numbers on a shared 4-CPU machine move with other jobs (docker ps showed no running containers). "gzip -6" is zlib's DEFLATE at gzip's default level.
import time, zlib
import cramjam, lz4.frame, zstandard
CODECS = { # name: (compress, decompress)
"gzip -6": (lambda d: zlib.compress(d, 6), zlib.decompress),
"snappy": (lambda d: bytes(cramjam.snappy.compress_raw(d)),
lambda c: bytes(cramjam.snappy.decompress_raw(c))),
"lz4": (lz4.frame.compress, lz4.frame.decompress),
**{f"zstd -{level}": (zstandard.ZstdCompressor(level=level).compress,
zstandard.ZstdDecompressor().decompress) for level in (1, 3, 19)},
}
def best_of_3(fn, arg):
times = []
for _ in range(3):
start = time.perf_counter()
result = fn(arg)
times.append(time.perf_counter() - start)
return result, min(times)
print("load average:", open("/proc/loadavg").read().split()[:3])
for path in ("data/orders.jsonl", "../arrow/orders.arrow"):
data = open(path, "rb").read()
print(f"{path} ({len(data) / 1e6:.1f} MB)",
" size ratio compress speed decompress speed")
base = None
for name, (compress, decompress) in CODECS.items():
packed, c_secs = best_of_3(compress, data)
restored, d_secs = best_of_3(decompress, packed)
assert restored == data
base = base or (c_secs, d_secs)
print(f" {name:9} {len(data) / len(packed):17.1f}x {base[0] / c_secs:12.2f}x gzip "
f"{base[1] / d_secs:13.2f}x gzip")
print("load average:", open("/proc/loadavg").read().split()[:3])load average: ['3.39', '2.60', '2.31'] data/orders.jsonl (23.6 MB) size ratio compress speed decompress speed gzip -6 14.0x 1.00x gzip 1.00x gzip snappy 6.5x 9.09x gzip 4.56x gzip lz4 7.2x 11.07x gzip 4.31x gzip zstd -1 11.4x 6.94x gzip 4.17x gzip zstd -3 11.0x 4.53x gzip 3.33x gzip zstd -19 20.6x 0.01x gzip 8.60x gzip ../arrow/orders.arrow (11.8 MB) size ratio compress speed decompress speed gzip -6 6.7x 1.00x gzip 1.00x gzip snappy 2.8x 56.00x gzip 4.06x gzip lz4 2.9x 54.50x gzip 4.86x gzip zstd -1 5.2x 22.56x gzip 1.69x gzip zstd -3 5.5x 11.68x gzip 2.29x gzip zstd -19 9.2x 0.08x gzip 1.98x gzip load average: ['7.50', '5.24', '3.67']

Size ratios are exact and identical on every run. Speed ratios are not: across four runs, zstd 126 -1 compressed the JSON 6.7 to 7.3 times faster than gzip and LZ4 decompressed it 3.3 to 5.9 times faster, and the load average climbed from 3.4 to 7.5 during the run shown as other jobs started. Read the speeds as bands; the Python wrappers' extra copies also understate the fastest codecs. Four lessons hold regardless:
Snappy 6,616 and LZ4 trade half the ratio for speed. They compressed 9 to 69 times faster than gzip but left files twice as large. Use them where CPU or latency matters more than storage.
zstd -1 is the best default. It kept about 80 percent of gzip's ratio while compressing 7 to 23 times faster and decompressing 1.7 to 4 times faster.
High levels are for write-once archives. zstd -19 shrank the JSON 20.6 times, 47 percent better than gzip, at a hundredth of its compression speed, yet still decompressed faster than gzip.
Layout beats codec. The best codec left the JSON at 1.1 MB; Parquet 129 's encodings plus zstd reached 0.8 MB (Compression inside Parquet) and Avro 129 with deflate 1.3 MB (Object Container Files), while also allowing column or block skipping.
zstd -3 came out slightly larger than zstd -1 on this JSON, a reminder that levels are heuristics, not guarantees: measure on your own data before standardizing on a setting.