Avro 129 writes the fields in schema order and nothing else: no names, no type markers, no field numbers. This script encodes order 1 (orders_io.py loads the JSON with Decimal money and datetime timestamps), then each field alone:
import io, json
from fastavro import parse_schema, schemaless_writer
from orders_io import orders
schema = parse_schema(json.load(open("order.avsc")))
order = next(orders())
buf = io.BytesIO()
schemaless_writer(buf, schema, order)
print(f"Avro body: {len(buf.getvalue())} bytes (the JSON line is 253)")
for field in schema["fields"]:
one = io.BytesIO()
schemaless_writer(one, field["type"], order[field["name"]])
print(f"{field['name']:12} {one.getvalue().hex(' ')}")Avro body: 42 bytes (the JSON line is 253) order_id 02 customer_id ac 27 order_ts e0 94 df f2 83 65 channel 00 status 12 64 65 6c 69 76 65 72 65 64 currency 06 55 53 44 items 04 06 02 04 09 60 0a 02 04 06 54 00 coupon 00 discount 02 00 total 04 0f b4
Every int and long is a zigzag varint: zigzag maps 0, -1, 1, -2 to 0, 1, 2, 3, and the varint stores 7 bits per byte, the high bit meaning "more follows". Order id 1 becomes 02; customer 2518 becomes 5036, bytes ac 27. Strings are a length then the data (12 is 9, then "delivered"). The enum and union write an index (00), the array a block count (04 is 2 items), the items and a 00 end marker, and the decimal total the integer 4020.
That is 42 bytes against 253 for JSON, and less than Protocol Buffers 59,573 ' 52 (proto3 Messages and Wire Types), which tags every field. The price is that the bytes mean nothing without the exact writer's schema, so Avro never ships data alone: container files embed the schema (Object Container Files), and Kafka 129 messages carry a registry's schema ID (Apache Kafka and Managed Cloud Kafka).