A scheduled pipeline starts on the clock ("at 02:00 UTC, process yesterday"). It is easy to reason about and to backfill, since each run owns one interval, but it guesses: if the export lands at 02:20, the run finds nothing. An event-driven pipeline starts when its input appears, such as a landed file, a published table or a queue message, at the cost of more moving parts and less predictable load. Sensors sit between the two.
| Trigger | Starts when | Good for | Watch out for |
|---|---|---|---|
| Schedule (cron) | The clock says so | Daily reports, backfills | Inputs that arrive late |
| Sensor | Schedule, then input appears | Late or irregular files | Slots held while waiting |
| Data-aware | An upstream dataset updates | Chains of pipelines | Hidden coupling between teams |
| Event | An external message arrives | Low latency | Bursts, ordering, duplicates |
Airflow 3 129 supports all four (Assets Replace Datasets, Event-Driven Scheduling and Sensors and Task Mapping). Event-driven is not streaming: each run is still a batch, started at a better moment. BookNest's pipeline uses a schedule, because finance closes the books by day.