Columnar Byte Layout

How Columnar Stores Lay Out Bytes

A columnar store writes all values of one column contiguously, usually in chunks of many thousands of rows so that a file still splits into independent pieces (Parquet 129 's row groups, Row Groups and Pages).

The same orders stored row by row and column by column
The same orders stored row by row and column by column

Values of one type sit together, so dictionary and run-length encodings (Parquet Encodings) and general compressors work better. Even plain zlib shows it on six order columns written both ways:

The same bytes, compressed in row-major and column-major orderJavaScript
import csv, zlib
cols = ["order_id", "customer_id", "order_ts", "channel", "status", "total"]
with open("csv-out/orders.csv", newline="", encoding="utf-8") as f:
    rows = [[r[c] for c in cols] for r in csv.DictReader(f)]
row_major = "\n".join(",".join(r) for r in rows).encode()
col_major = "\n".join(",".join(r[i] for r in rows) for i in range(len(cols))).encode()
for name, blob in (("row-major", row_major), ("column-major", col_major)):
    packed = len(zlib.compress(blob, 6))
    print(f"{name:12} {len(blob):>10,} bytes -> zlib {packed:>9,} ({len(blob) / packed:.1f}x)")
Output
row-major     5,249,279 bytes -> zlib 1,179,871 (4.4x)
column-major  5,249,279 bytes -> zlib   983,846 (5.3x)

Same bytes, different order: the column-major version compresses 17 percent smaller. The price is paid on writes, since one new order touches every column, so columnar files are written in large, effectively immutable batches.