The OpenLineage Specification

Where did those 900 orders come from, and what did they reach? Lineage answers both, and OpenLineage (https://github.com/OpenLineage/OpenLineage 2,687 ) (Apache-2.0, an LF AI & Data project, 1.53.0 on 1 September 2026) is its open standard. A job (a recurring process in a namespace) has runs (one execution each, a UUID), which read and write datasets named by a system namespace (kafka://l1-kafka:9092, s3://warehouse) and a name within it. Producers send run events (START, RUNNING, COMPLETE, FAIL, ABORT) as JSON over HTTP, Kafka 129 , a file or the console, and facets add typed metadata: schemas, errors, SQL, column-level lineage, or symlinks to another name for a dataset.

Integrations emit the events for you: a Spark 129 listener, Airflow 129 's OpenLineage provider (Orchestration and Pipelines), the dbt-ol wrapper, Flink 129 and Trino 403,499 . The Spark one is a jar plus spark.extraListeners, but it fell short here: the 1.53.0 jar carries version-specific code up to Spark 4.0 (packages spark31 to spark40), and on Spark 4.1.3 the daily_sales job of Spark as a Lakehouse Engine (governance/ol_spark.py) sent run events such as booknest_lakehouse.replace_data with empty input and output lists. Where no integration works, the pipeline reports its own runs with the Python client.