All four codecs belong to the LZ77 family: they replace repeated byte sequences with back-references to earlier occurrences. They differ in how hard they search for matches and how they encode the result.
| Codec | Origin | Method | Levels | Python package |
|---|---|---|---|---|
| gzip (DEFLATE) | 1992; RFC 1951-1952 (1996) | LZ77 + Huffman coding | 1-9 | zlib (standard library) |
| Snappy 6,616 | Google, open-sourced 2011 | Byte-aligned LZ77, no entropy coding | none | cramjam (MIT) |
| LZ4 | Yann Collet, 2011 | Byte-aligned LZ77, tuned for speed | fast and HC modes | lz4 (BSD) |
| Zstandard 126 | Yann Collet at Facebook, 2016; RFC 8878 | LZ77 + Huffman + finite-state entropy | negative to 22 | zstandard (BSD) |
Gzip is everywhere: HTTP, .gz files, Avro 129 's required deflate codec. Its Huffman stage squeezes well but decompresses slowly. Snappy and LZ4 skip entropy coding to run at hundreds of megabytes per second per core; Snappy became the default of Hadoop-era tools, Parquet 129 writers and Spark 129 , and LZ4 suits latency-sensitive paths such as Kafka 129 producers (Apache Kafka and Managed Cloud Kafka). Zstandard (zstd) spans the whole range, from fast low levels to high levels that out-compress gzip, and decompresses quickly at every level. It also supports trained dictionaries for small messages such as single JSON events. Brotli (Google, RFC 7932) compresses text very well but slowly and lives mostly on the web. From the command line, zstd -T0 and pigz compress on all cores.