Erasure Coding

MinIO's Erasure-Coded Architecture

MinIO 30,943 stores objects as plain files on ordinary drives, with no metadata database. Drives form erasure sets of 2 to 16; SipHash of the object name picks the set, and each object part is cut into data shards plus Reed-Solomon parity shards, one per drive. With EC:2 on four drives, any two shards rebuild the data. A per-object xl.meta on every drive records the version, the layout and a HighwayHash bitrot checksum per shard, so silent corruption is detected on read and repaired from the other shards.

One multipart object on a four-drive erasure set with EC:2 (parity placement varies per object)
One multipart object on a four-drive erasure set with EC:2 (parity placement varies per object)

The multipart object from Consistency and Versioning shows the layout, and survives losing two drives' shards:

erasure.sh: shards on four drives, two of them deleted, the object intactShell
# Where MinIO put orders.jsonl: shards of every part on each of the four drives.
docker exec l1-minio sh -c 'cd /data; for d in disk*; do
  echo "$d: $(du -b $d/booknest-demo/orders.jsonl/*/part.* | sed "s|\t.*/| |" | tr "\n" " ")"
done'
# Lose two of the four drives' shards: the object still reads back byte for byte.
docker exec l1-minio sh -c 'rm -rf /data/disk[13]/booknest-demo/orders.jsonl'
mc cat lake/booknest-demo/orders.jsonl | sha256sum | cut -c1-16
sha256sum "/mnt/d/Books/Data Engineering/demos/ch03/data/orders.jsonl" | cut -c1-16
mc admin heal -r lake/booknest-demo >/dev/null 2>&1   # rebuild the missing shards now
Output
disk1: 8389120 part.1 3399323 part.2
disk2: 8389120 part.1 3399323 part.2
disk3: 8389120 part.1 3399323 part.2
disk4: 8389120 part.1 3399323 part.2
ff51790e08776738
ff51790e08776738

Each drive holds 11,788,443 bytes, so the 23.6 MB object occupies 47.2 MB, the 2x of EC:2 plus bitrot hashes. With two drives' shards gone, the read still matched JSON, Columnar and Binary Formats's SHA-256; mc admin heal then rewrites them. In production the drives sit in four servers; here they share one volume.