Project Nessie

Project Nessie and Git-Like Table Versioning

Iceberg 129 versions each table on its own. Project Nessie (https://github.com/projectnessie/nessie 1,517 ) (Apache-2.0, started at Dremio 249,796 , 0.108.8 on 9 September 2026) versions the whole catalog: each commit records every table's pointer, so a branch holds a consistent state of many tables, and a merge publishes them atomically. Branches, tags and merges work as in Git 1,932 (Package Managers and DevOps), over a version store in a JDBC database, RocksDB, MongoDB 1,815 , DynamoDB 24 or memory, among others. Nessie serves Iceberg REST at /iceberg/<branch>, so an engine picks a branch through its catalog URI.

Its image would not download here (eight TLS failures from ghcr.io), so demos/ch08/catalogs/run_nessie.sh runs the 276 MB server jar with Java 21, an in-memory store and warehouse s3://warehouse/nessie/. The script copies the books table onto main, branches, cuts a price, and merges:

nessie_branch.py: branch the catalog, change a table there, compare, mergePython
"""Git-like versioning with Nessie: branch the catalog, change a table there, merge."""
import json, urllib.request
from lake import spark
API = "http://localhost:31183/api/v2"
def nessie(path, body=None):                    # a tiny client for Nessie's REST API v2
    req = urllib.request.Request(API + path, json.dumps(body).encode() if body else None,
                                 {"Content-Type": "application/json"})
    return json.load(urllib.request.urlopen(req))
def catalog(name, ref):                         # one Spark catalog per Nessie branch
    c = f"spark.sql.catalog.{name}"
    spark.conf.set(c, "org.apache.iceberg.spark.SparkCatalog")
    spark.conf.set(c + ".uri", f"http://localhost:31183/iceberg/{ref}")
    for k in ("type", "io-impl", "s3.endpoint", "s3.path-style-access", "s3.access-key-id",
              "s3.secret-access-key", "client.region"):     # as in lake.py: MinIO's keys
        spark.conf.set(f"{c}.{k}", spark.conf.get(f"spark.sql.catalog.lake.{k}"))
catalog("nm", "main")
spark.sql("CREATE NAMESPACE IF NOT EXISTS nm.booknest")
spark.sql("""CREATE TABLE nm.booknest.books (book_id INT, title STRING, author STRING,
             genre STRING, price DECIMAL(9,2), year INT) USING iceberg""")
spark.sql("INSERT INTO nm.booknest.books SELECT * FROM lake.booknest.books")
nessie("/trees?name=spring-sale&type=BRANCH", nessie("/trees/main")["reference"])
catalog("sale", "spring-sale")
spark.sql("UPDATE sale.booknest.books SET price = price - 5 WHERE genre = 'Technology'")
prices = lambda: print(spark.sql("""SELECT m.price AS main, s.price AS branch
    FROM nm.booknest.books m JOIN sale.booknest.books s USING (book_id)
    WHERE book_id = 2""").first())                            # Patterns of the Deep Web
prices()
for e in nessie("/trees/spring-sale/history")["logEntries"]:
    print(e["commitMeta"]["hash"][:8], e["commitMeta"]["message"])
main, sale = (nessie(f"/trees/{b}")["reference"]["hash"] for b in ("main", "spring-sale"))
r = nessie(f"/trees/main@{main}/history/merge",          # merge the branch into main
           {"fromRefName": "spring-sale", "fromHash": sale})
print("merged; main is now at", r["resultantTargetHash"][:8])
spark.sql("REFRESH TABLE nm.booknest.books")
prices()
Output
Row(main=Decimal('39.50'), branch=Decimal('34.50'))
a75bb48e Update ICEBERG_TABLE booknest.books
5e5e401a Update ICEBERG_TABLE booknest.books
f011e93a Create ICEBERG_TABLE booknest.books
811158fb update namespace booknest
merged; main is now at 5ddc2237
Row(main=Decimal('34.50'), branch=Decimal('34.50'))

Readers of main kept the old price of Patterns of the Deep Web until the merge, and the branch's log lists each Iceberg commit (the insert, then the update) as a Nessie commit. Loading into a branch, checking it and merging only on success is write-audit-publish (Branches and Tags), here across many tables.