ORC cuts rows into stripes (its row groups) and stores each column of a stripe as streams: PRESENT bits for nulls, DATA, LENGTH for strings and lists, DICTIONARY_DATA for dictionaries. A stripe opens with row-index streams holding statistics per 10,000 rows and closes with a footer listing its streams; the file ends with a footer (schema, stripe locations, statistics), a postscript and a final byte giving the postscript's length. This script writes the Parquet 129 orders as ORC and reads the layout with Java orc-tools 2.3.1 (an uber jar on Maven 129 Central):
python orc_write.py
java -Duser.timezone=UTC -jar "$ORC_TOOLS" meta orders.orc > meta.txt 2> /dev/null
grep -E "Stripe: offset" meta.txt
grep -E "column 5 section [DL]|Encoding column 5" meta.txt | head -4 # status, stripe 1
grep -E "^File length" meta.txt100,000 rows in 4 stripes, ZSTD, format 0.12, writer ORC C++ 2.2.2
orders.orc 699,869 bytes
../parquet/orders.parquet 798,714 bytes
Stripe: offset: 3 data: 225249 rows: 32768 tail: 270 index: 1510
Stripe: offset: 227032 data: 224471 rows: 32768 tail: 266 index: 1524
Stripe: offset: 453293 data: 229003 rows: 32768 tail: 270 index: 1544
Stripe: offset: 684110 data: 13627 rows: 1696 tail: 253 index: 484
Stream: column 5 section DATA start: 119080 length 4911
Stream: column 5 section DICTIONARY_DATA start: 123991 length 29
Stream: column 5 section LENGTH start: 124020 length 7
Encoding column 5: DICTIONARY_V2[3]
File length: 699869 bytesColumn 5 is status (the root struct is 0): a 29-byte dictionary of the stripe's three values, their lengths, and 4,911 bytes of indexes for 32,768 rows. orc_write.py passes dictionary_key_size_threshold=0.8 because the C++ writer's default of 0 disables dictionaries.