Measuring Size and Speed

Measuring Size and Speed on BookNest Data

Vendor benchmarks use their own data; this one uses two of this chapter's files. orders.jsonl is 23.6 MB of row-oriented text, and orders.arrow (Arrow IPC and Feather Files) is 11.8 MB of uncompressed, decoded columns. Each codec compresses and decompresses each file three times, keeps the fastest run, checks the round trip, and reports speeds relative to gzip, since absolute numbers on a shared 4-CPU machine move with other jobs (docker ps showed no running containers). "gzip -6" is zlib's DEFLATE at gzip's default level.

bench.py: size and speed of five codecs on two BookNest filesPython
import time, zlib
import cramjam, lz4.frame, zstandard
CODECS = {                                     # name: (compress, decompress)
    "gzip -6": (lambda d: zlib.compress(d, 6), zlib.decompress),
    "snappy": (lambda d: bytes(cramjam.snappy.compress_raw(d)),
               lambda c: bytes(cramjam.snappy.decompress_raw(c))),
    "lz4": (lz4.frame.compress, lz4.frame.decompress),
    **{f"zstd -{level}": (zstandard.ZstdCompressor(level=level).compress,
                          zstandard.ZstdDecompressor().decompress) for level in (1, 3, 19)},
}
def best_of_3(fn, arg):
    times = []
    for _ in range(3):
        start = time.perf_counter()
        result = fn(arg)
        times.append(time.perf_counter() - start)
    return result, min(times)
print("load average:", open("/proc/loadavg").read().split()[:3])
for path in ("data/orders.jsonl", "../arrow/orders.arrow"):
    data = open(path, "rb").read()
    print(f"{path} ({len(data) / 1e6:.1f} MB)",
          "  size ratio   compress speed   decompress speed")
    base = None
    for name, (compress, decompress) in CODECS.items():
        packed, c_secs = best_of_3(compress, data)
        restored, d_secs = best_of_3(decompress, packed)
        assert restored == data
        base = base or (c_secs, d_secs)
        print(f"  {name:9} {len(data) / len(packed):17.1f}x {base[0] / c_secs:12.2f}x gzip "
              f"{base[1] / d_secs:13.2f}x gzip")
print("load average:", open("/proc/loadavg").read().split()[:3])
Output
load average: ['3.39', '2.60', '2.31']
data/orders.jsonl (23.6 MB)   size ratio   compress speed   decompress speed
  gzip -6                14.0x         1.00x gzip          1.00x gzip
  snappy                  6.5x         9.09x gzip          4.56x gzip
  lz4                     7.2x        11.07x gzip          4.31x gzip
  zstd -1                11.4x         6.94x gzip          4.17x gzip
  zstd -3                11.0x         4.53x gzip          3.33x gzip
  zstd -19               20.6x         0.01x gzip          8.60x gzip
../arrow/orders.arrow (11.8 MB)   size ratio   compress speed   decompress speed
  gzip -6                 6.7x         1.00x gzip          1.00x gzip
  snappy                  2.8x        56.00x gzip          4.06x gzip
  lz4                     2.9x        54.50x gzip          4.86x gzip
  zstd -1                 5.2x        22.56x gzip          1.69x gzip
  zstd -3                 5.5x        11.68x gzip          2.29x gzip
  zstd -19                9.2x         0.08x gzip          1.98x gzip
load average: ['7.50', '5.24', '3.67']
Compression size ratios on two BookNest files (higher means smaller output)
Compression size ratios on two BookNest files (higher means smaller output)

Size ratios are exact and identical on every run. Speed ratios are not: across four runs, zstd 126 -1 compressed the JSON 6.7 to 7.3 times faster than gzip and LZ4 decompressed it 3.3 to 5.9 times faster, and the load average climbed from 3.4 to 7.5 during the run shown as other jobs started. Read the speeds as bands; the Python wrappers' extra copies also understate the fastest codecs. Four lessons hold regardless:

zstd -3 came out slightly larger than zstd -1 on this JSON, a reminder that levels are heuristics, not guarantees: measure on your own data before standardizing on a setting.