Encoding turns a page's values into bytes before any compression, and the right one depends on the column:
PLAIN: values back to back, fixed width or length-prefixed.
RLE_DICTIONARY: a dictionary page of distinct values, then indexes in the RLE/bit-packing hybrid (which also stores levels and booleans). Writers fall back to PLAIN when the dictionary passes 1 MB in Arrow 129 .
DELTA_BINARY_PACKED: bit-packed differences between consecutive integers, for sorted ids and timestamps; DELTA_LENGTH_BYTE_ARRAY and DELTA_BYTE_ARRAY apply the idea to string lengths and prefixes.
BYTE_STREAM_SPLIT: regroups fixed-width values byte by byte for a compressor, meant for floats; specification 2.14.0 adds ALP for floats as a preview.
This experiment writes five columns alone, uncompressed, with each encoding (n/a: pyarrow 129 refuses the pairing):
python encodings.py # each column alone, written uncompressed with each encodingKB before compression PLAIN RLE_DICTIONARY DELTA_BINARY_PACKED BYTE_STREAM_SPLIT order_id 800 1,003 2 800 order_ts 800 1,002 272 800 customer_id 400 183 176 400 status 1,293 23 n/a n/a total 400 121 188 400
The default, dictionary everywhere, is the worst choice for the unique, sorted order_id: 1,003 KB against 2 KB with delta encoding, which stores a start value and a run of differences of 1. status shrinks 56-fold with a six-string dictionary. The writer options encode these findings:
"""Parquet writer settings for BookNest orders, chosen from the encoding comparison."""
import pyarrow.parquet as pq
DELTA = ["order_id", "order_ts"] # sorted integers: store differences
DICTIONARY = ["customer_id", "channel", "status", "currency", "coupon", "discount", "total",
"items.list.element.book_id", "items.list.element.qty",
"items.list.element.unit_price"]
OPTIONS = dict(row_group_size=25_000, use_dictionary=DICTIONARY,
column_encoding={c: "DELTA_BINARY_PACKED" for c in DELTA},
store_decimal_as_integer=True, # DECIMAL(9,2) as INT32, not byte arrays
write_page_index=True, sorting_columns=[pq.SortingColumn(0)])With pyarrow's defaults (dictionary everywhere, decimals as fixed-length byte arrays) the zstd 126 file was 1,316,571 bytes; these options make it 798,714, 39 percent smaller. Other writers have their own heuristics, so check what they did (Parquet Metadata in Practice) before tuning.