What to Test in a Pipeline

Test code before it ships and data every time it moves. For BookNest's daily pipeline:

Test layers for a data pipeline
Layer Question BookNest example Runs
Unit Does a function transform correctly? extract_orders keeps one UTC day CI (Unit Testing DAGs with pytest)
Integration Does a step work against real systems? load_orders twice gives one copy CI, disposable PostgreSQL 1,289
DAG integrity Does the DAG parse with the right shape? No import errors, publish after check CI (Unit Testing DAGs with pytest)
Data tests Is the transformed data valid? dbt 37,942 not_null, relationships Every run (4.11.7)
Reconciliation Do totals match the source? check_day: lines and gross Every run (BookNest's Daily Pipeline)
Contract Does the producer still send what we expect? ODCS checks on the export Every run (Data Contracts)

Two rules carry most of the value. First, make tests deterministic: fix the day (ds), the time zone and the random seed, as BookNest's sample data does. Second, put the run-time checks before the step that publishes, and make them block it: a check that only logs a warning after the dashboard refreshed is a post-mortem, not a test. Idempotent steps (Idempotency and Safe Backfills) are what make re-running after a fixed failure safe.