Each loop, the scheduler creates DAG runs that are due, finds task instances whose upstream tasks are done, and, in a short critical section, queues them for the executor within parallelism, pool and per-DAG limits. Several schedulers can run for high availability on PostgreSQL 12 1,289 + or MySQL 8.0 524 +; they coordinate through row locks (SELECT ... FOR UPDATE with SKIP LOCKED and NOWAIT), not leader election.
The scheduler never reads your Python files. The DAG processor, a separate process that Airflow 3 129 requires, lists each DAG bundle (here the dags/ folder; Git 1,932 bundles work too) every refresh_interval seconds, re-parses each file every min_file_process_interval seconds in a child process, and stores the result as a serialized DAG, JSON that the scheduler and UI use. Its log reports each parse (columns trimmed):
Bundle File Path PID Current Duration # DAGs # Errors Last Duration ----------- ----------------- ----- ---------------- ------ -------- ------------- dags-folder booknest_daily.py 1 0 0.27s dags-folder arch_demo.py 1 0 0.26s
With this stack's defaults (300 s and 30 s), a new file can take five minutes to appear, and top-level code runs on every parse: a query outside a task runs every 30 seconds whether or not the DAG runs. Keep module-level code to imports and DAG definitions. Because the scheduler never executes user code, a DAG file that hangs slows only the DAG processor.