Parquet Encodings

Encoding turns a page's values into bytes before any compression, and the right one depends on the column:

This experiment writes five columns alone, uncompressed, with each encoding (n/a: pyarrow 129 refuses the pairing):

encodings.sh: one column, four encodingsShell
python encodings.py    # each column alone, written uncompressed with each encoding
Output
KB before compression  PLAIN  RLE_DICTIONARY  DELTA_BINARY_PACKED  BYTE_STREAM_SPLIT
order_id                800           1,003                    2                800
order_ts                800           1,002                  272                800
customer_id             400             183                  176                400
status                1,293              23                  n/a                n/a
total                   400             121                  188                400

The default, dictionary everywhere, is the worst choice for the unique, sorted order_id: 1,003 KB against 2 KB with delta encoding, which stores a start value and a run of differences of 1. status shrinks 56-fold with a six-string dictionary. The writer options encode these findings:

options.py: per-column encodings for the orders filePython
"""Parquet writer settings for BookNest orders, chosen from the encoding comparison."""
import pyarrow.parquet as pq
DELTA = ["order_id", "order_ts"]                       # sorted integers: store differences
DICTIONARY = ["customer_id", "channel", "status", "currency", "coupon", "discount", "total",
              "items.list.element.book_id", "items.list.element.qty",
              "items.list.element.unit_price"]
OPTIONS = dict(row_group_size=25_000, use_dictionary=DICTIONARY,
               column_encoding={c: "DELTA_BINARY_PACKED" for c in DELTA},
               store_decimal_as_integer=True,          # DECIMAL(9,2) as INT32, not byte arrays
               write_page_index=True, sorting_columns=[pq.SortingColumn(0)])

With pyarrow's defaults (dictionary everywhere, decimals as fixed-length byte arrays) the zstd 126 file was 1,316,571 bytes; these options make it 798,714, 39 percent smaller. Other writers have their own heuristics, so check what they did (Parquet Metadata in Practice) before tuning.