Nested Data Levels

Nested Data and Repetition and Definition Levels

Parquet 129 stores only flat leaf columns, so a nested field such as an order's items list is shredded: each leaf (items.list.element.book_id and its siblings) becomes its own column, and two small integers per value let a reader rebuild the structure, as the Dremel paper describes. The definition level says how many optional or repeated ancestors of a value are actually present, so it encodes nulls and empty lists; the repetition level says at which repeated ancestor a new value starts, where 0 means a new record. The schema fixes each column's maximum levels:

levels.sh: maximum levels for flat, nested and optional columnsShell
python levels.py | grep -E "items|coupon|^order_id"
Output
order_id                       INT64      max_rep=0 max_def=0
items.list.element.book_id     INT32      max_rep=1 max_def=1
items.list.element.qty         INT32      max_rep=1 max_def=1
items.list.element.unit_price  INT32      max_rep=1 max_def=1
coupon                         BYTE_ARRAY max_rep=0 max_def=1

levels.py prints each column's maximum levels from pyarrow 129 's ParquetFile(...).schema; a required flat column needs none. These are the levels for orders 1, 2 and 16 (the first with a coupon):

Levels that shred three orders' items and coupons into flat columns
Order book_id values Repetition levels Definition levels coupon (definition level)
1 3, 5 0, 1 1, 1 null (0)
2 3, 6, 5 0, 1, 1 1, 1, 1 null (0)
16 1 0 1 READMORE15 (1)

Repetition level 1 means "another element of the same list" and 0 starts the next order; definition level 0 on coupon marks a null, so no value is stored. With the list and its elements required, only the repeated list counts toward book_id's definition level, and an empty list would be one entry at level 0. With Arrow 129 's default (everything nullable) the maximum would be 4. The levels are RLE-encoded and nearly free for regular data, and a query on book_id reads only that leaf column.