Scheduled vs Event-Driven

Scheduling Versus Event-Driven Pipelines

A scheduled pipeline starts on the clock ("at 02:00 UTC, process yesterday"). It is easy to reason about and to backfill, since each run owns one interval, but it guesses: if the export lands at 02:20, the run finds nothing. An event-driven pipeline starts when its input appears, such as a landed file, a published table or a queue message, at the cost of more moving parts and less predictable load. Sensors sit between the two.

How pipeline runs are triggered
Trigger Starts when Good for Watch out for
Schedule (cron) The clock says so Daily reports, backfills Inputs that arrive late
Sensor Schedule, then input appears Late or irregular files Slots held while waiting
Data-aware An upstream dataset updates Chains of pipelines Hidden coupling between teams
Event An external message arrives Low latency Bursts, ordering, duplicates

Airflow 3 129 supports all four (Assets Replace Datasets, Event-Driven Scheduling and Sensors and Task Mapping). Event-driven is not streaming: each run is still a batch, started at a better moment. BookNest's pipeline uses a schedule, because finance closes the books by day.