Why a Data Quality Layer

Why Data Quality Needs a Layer of Its Own

A table format makes a commit atomic, not its rows right: Streaming into the Lakehouse's sink atomically committed 1,800 events for orders the orders table has never heard of. Bad data rarely breaks a pipeline; it flows into dashboards until someone notices a wrong number, so checks belong to the platform and run on every load.

Checks fall into recurring dimensions: completeness (no missing customer on a placed order), uniqueness (one row per event), validity (statuses from a fixed set), consistency (every event's order exists; the gold table sums to its source), freshness and volume (Anomaly Detection). Where a check runs matters too: at ingestion it can quarantine a batch, as Iceberg 129 's write-audit-publish branch does (Branches and Tags); after a transformation it guards a model; on a schedule it catches the rest. A failure should stop the pipeline (Orchestration and Pipelines) or alert, never just write a report nobody opens.