Packfiles

Packfiles, Delta Compression and Loose Objects

Objects stored one per file are loose: quick to write, but each fills at least one disk block and every version is stored whole. Commit BookNest's six-book seed.json catalog, then 20 small edits to it, and watch git gc:

Creating 20 versions of a 1.6 KB file, packing them and inspecting the packShell
git commit -qam "Mention the runtime and database"
cp ../booknest/db/seed.json db/
git add db/seed.json && git commit -qm "Add the six-book seed catalog"
for i in $(seq 1 20); do
  sed -i "s/\"rating\": [0-9.]*/\"rating\": 4.$i/" db/seed.json
  git commit -qam "Adjust ratings, round $i"
done
git count-objects -vH
git gc
git count-objects -vH
git verify-pack -v .git/objects/pack/pack-*.idx | grep blob | head -8 | cut -c1-12,41-
Output
count: 96
size: 384.00 KiB
in-pack: 0
...
count: 0
size: 0 bytes
in-pack: 96
packs: 2
size-pack: 14.03 KiB
...
f8dabc81b9b9 blob   1653 702 4028
...
b1a84f3b3b80 blob   46 58 5383 1 f8dabc81b9b967906ff7559e54f49325d0c380db
06a7a1f18342 blob   46 51 5600 1 f8dabc81b9b967906ff7559e54f49325d0c380db
25f0195fb86b blob   34 49 5810 1 f8dabc81b9b967906ff7559e54f49325d0c380db

Ninety-six loose objects in 384 KiB of disk blocks became 14 KiB of packs. A packfile concatenates compressed objects, and its .idx maps each hash to an offset. The big saving comes from deltas: verify-pack lists hash (cut to 12 digits here), type, size, size in the pack and offset, then for deltas the chain depth and base. The newest seed.json is stored whole (1,653 bytes, 702 compressed) and older versions as deltas of 34 to 46 bytes: Git 1,932 keeps the version you read most in full and picks delta bases by similarity. The same format travels over the network (Git's Transport Protocols). The second, tiny pack is a cruft pack holding the README version you staged and then overwrote in The Index, with an .mtimes file recording when it was last seen.