Arrow Memory Format

The Arrow Columnar Memory Format

An Arrow 129 array is a few contiguous buffers: a validity bitmap (one bit per value, 1 meaning present), then type-specific buffers. Fixed-width types have one data buffer; strings and lists add an offsets buffer whose entries mark where each value starts and ends; dictionary arrays hold integer indices into a dictionary array. Buffers are recommended to be 64-byte aligned to match SIMD register width, and no value needs parsing:

buffers.py: the buffers behind a string column with nullsPython
import pyarrow as pa
coupons = pa.array(["SPRING10", None, "READMORE15", None])   # four orders' coupons
validity, offsets, data = coupons.buffers()
print("type:", coupons.type, "| length:", len(coupons), "| nulls:", coupons.null_count)
print("validity bitmap:", f"{validity.to_pybytes()[0]:08b}", "(read right to left)")
print("offsets (int32):", pa.Array.from_buffers(pa.int32(), 5, [None, offsets]).to_pylist())
print("data bytes:     ", data.to_pybytes())
channel = pa.array(["ios", "ios", "web", "ios"]).dictionary_encode()
print("dictionary:", channel.dictionary.to_pylist(), "indices:", channel.indices.to_pylist())
Output
type: string | length: 4 | nulls: 2
validity bitmap: 00000101 (read right to left)
offsets (int32): [0, 8, 8, 18, 18]
data bytes:      b'SPRING10READMORE15'
dictionary: ['ios', 'web'] indices: [0, 0, 1, 0]
The three buffers of the coupons array
The three buffers of the coupons array

Null values occupy a bit and an empty offset range, not data. Because the layout is fully specified, a library in another language, or another process sharing the memory, can use these buffers as they are: reading Arrow data needs no deserialization, only pointer arithmetic. Columnar format 1.4 added binary views (adapted from TU Munich's Umbra database, with short strings stored inline) and list views, whose offsets need not be in order.