Row Groups and Pages

Row Groups, Column Chunks and Pages

A Parquet 129 file is a hierarchy. Rows are cut horizontally into row groups; within a row group each column is stored contiguously as a column chunk; each chunk is a sequence of pages, the unit of encoding and compression (an optional dictionary page, then data pages of about 1 MB by default). The footer at the end, a Thrift compact-protocol FileMetaData struct, holds the schema and every chunk's offset, encodings, codec and statistics, so a reader fetches the footer first and then only the byte ranges it needs.

The layout of orders.parquet: row groups, column chunks, pages and the footer
The layout of orders.parquet: row groups, column chunks, pages and the footer

write_parquet.py writes the 100,000 orders, with exact decimal money and every field but coupon required, as a nested file of four 25,000-row groups (options.py is Parquet Encodings's). Later chapters reuse it as demos/ch03/out/orders.parquet.

write_parquet.py: the BookNest orders as ParquetPython
import json, struct
from datetime import datetime
from decimal import Decimal
import pyarrow as pa, pyarrow.parquet as pq
from options import OPTIONS
money = pa.decimal128(9, 2)
item = pa.struct([pa.field("book_id", pa.int32(), False), pa.field("qty", pa.int8(), False),
                  pa.field("unit_price", money, False)])
schema = pa.schema([
    pa.field("order_id", pa.int64(), False), pa.field("customer_id", pa.int32(), False),
    pa.field("order_ts", pa.timestamp("ms", tz="UTC"), False),
    pa.field("channel", pa.string(), False), pa.field("status", pa.string(), False),
    pa.field("currency", pa.string(), False),
    pa.field("items", pa.list_(pa.field("element", item, False)), False),
    pa.field("coupon", pa.string()), pa.field("discount", money, False),
    pa.field("total", money, False)])
rows = []
for line in open("data/orders.jsonl"):
    o = json.loads(line, parse_float=Decimal)
    o["order_ts"] = datetime.fromisoformat(o["order_ts"])
    rows.append(o)
table = pa.Table.from_pylist(rows, schema=schema)
table = table.replace_schema_metadata({"booknest.source": "orders.jsonl, seed 7"})
pq.write_table(table, "orders.parquet", compression="zstd", **OPTIONS)
with open("orders.parquet", "rb") as f:
    head = f.read(4)
    f.seek(-8, 2)                                  # footer length + magic close the file
    footer_len, tail = struct.unpack("<I4s", f.read(8))
    print(head, tail, f"{f.tell():,} bytes, footer {footer_len:,} bytes")
Output
b'PAR1' b'PAR1' 798,714 bytes, footer 7,155 bytes

The 23.6 MB of JSON Lines became 0.8 MB, with a 7 KB footer describing 4 row groups of 12 column chunks. Row groups are the unit of parallelism and skipping: a Spark 129 task or DuckDB 61,228 thread takes whole groups. Arrow 129 's default of 1,048,576 rows per group would have made one group here; the Parquet documentation suggests 512 MB to 1 GB groups for large data. Thousands of tiny files and groups are the classic data-lake slowdown that table compaction (Lakehouses, Data Quality and Governance) cures.