Scheduler and DAG Processor

The Scheduler and the DAG Processor

Each loop, the scheduler creates DAG runs that are due, finds task instances whose upstream tasks are done, and, in a short critical section, queues them for the executor within parallelism, pool and per-DAG limits. Several schedulers can run for high availability on PostgreSQL 12 1,289 + or MySQL 8.0 524 +; they coordinate through row locks (SELECT ... FOR UPDATE with SKIP LOCKED and NOWAIT), not leader election.

The scheduler never reads your Python files. The DAG processor, a separate process that Airflow 3 129 requires, lists each DAG bundle (here the dags/ folder; Git 1,932 bundles work too) every refresh_interval seconds, re-parses each file every min_file_process_interval seconds in a child process, and stores the result as a serialized DAG, JSON that the scheduler and UI use. Its log reports each parse (columns trimmed):

Output of 3
Bundle       File Path          PID    Current Duration  # DAGs  # Errors  Last Duration
-----------  -----------------  -----  ----------------  ------  --------  -------------
dags-folder  booknest_daily.py                                1         0  0.27s
dags-folder  arch_demo.py                                     1         0  0.26s

With this stack's defaults (300 s and 30 s), a new file can take five minutes to appear, and top-level code runs on every parse: a query outside a task runs every 30 seconds whether or not the DAG runs. Keep module-level code to imports and DAG definitions. Because the scheduler never executes user code, a DAG file that hangs slows only the DAG processor.