Many pipeline failures start upstream: an app team renames a field, and the nightly job breaks or loads nonsense. A data contract is the producer's written, versioned promise about a dataset: schema and types, quality rules, freshness and an owner. The consumer checks each delivery against it at the boundary, before loading. The Open Data Contract Standard (ODCS v3.2.0, Apache-2.0, from the Bitol project at the Linux Foundation's LF AI & Data) gives contracts a common YAML shape:
apiVersion: v3.2.0
kind: DataContract
id: booknest-orders
version: 1.0.0
status: active
schema:
- name: orders
properties:
- {name: order_id, logicalType: integer, required: true, primaryKey: true}
- {name: order_ts, logicalType: timestamp, required: true}
- {name: status, logicalType: string, required: true}
- {name: total, logicalType: number, required: true}
quality:
- {type: library, metric: rowCount, mustBeGreaterThan: 0, severity: error}
slaProperties: [{property: frequency, value: 1, unit: d}]A 30-line checker, demos/ch07/contract/check_contract.py, tests a good day, then a release that renamed total and switched order_ts to local text:
$ day=landing/orders/date=2026-06-28/orders.jsonl $ python check_contract.py orders.odcs.yaml $day; echo "exit $?" booknest-orders v1.0.0: 186 rows, 0 violations exit 0 $ jq -c '.amount = .total | del(.total) | .order_ts = "28/06/2026 09:15"' $day > broken.jsonl $ python check_contract.py orders.odcs.yaml broken.jsonl; echo "exit $?" booknest-orders v1.0.0: 186 rows, 372 violations row 1: order_ts='28/06/2026 09:15' is not a timestamp ... exit 1
An orchestrator acts on the non-zero exit: the task fails, nothing downstream runs, and the alert names the broken promise, not a SQL error three steps later. Ideally the producer runs the same check in its own CI. Anomalies, Contracts, Lineage returns to ODCS and enforces a contract on BookNest's order events.