An object container file (.avro) is Avro 129 's self-describing form: the magic bytes Obj plus byte 1, a metadata map holding the schema and codec, and a random 16-byte sync marker. Blocks of records follow, each compressed as a unit and closed by the marker, so a reader starting mid-file, such as one Spark 129 task, can find the next block boundary.
import json, os
from fastavro import block_reader, parse_schema, writer
from orders_io import orders
with open("orders.avro", "wb") as out:
writer(out, parse_schema(json.load(open("order.avsc"))), orders(), codec="deflate")
with open("orders.avro", "rb") as f:
print("magic:", f.read(4))
f.seek(0)
reader = block_reader(f)
counts = [block.num_records for block in reader]
print("metadata:", sorted(reader.metadata), "| sync:", reader._header["sync"].hex())
size = os.path.getsize("orders.avro")
print(f"{len(counts)} blocks, {sum(counts):,} records, {size:,} bytes")magic: b'Obj\x01' metadata: ['avro.codec', 'avro.schema'] | sync: 75a1ed6d75dc3a452d60833c43f081c9 260 blocks, 100,000 records, 1,316,812 bytes
The deflate file is 1.32 MB, 18 times smaller than the JSON Lines. Every implementation must read the null and deflate codecs; snappy, zstandard, bzip2 and xz are optional, so deflate is the safe choice for unknown readers (Compression Codecs Compared compares codecs). Logical-type support varies: DuckDB 1.5.5 61,228 's read_avro shows the decimals as BLOB, while Spark (Batch Processing with Apache Spark) reads DECIMAL(9,2). To see what an unknown file holds, run fastavro --schema orders.avro.