Why Orchestrate dbt

Why Orchestrate dbt Instead of Shelling Out

dbt 37,942 build --select fct_sales in a BashOperator is one opaque task. If a test fails, Airflow 129 knows only that the task failed; a retry rebuilds every selected model; the log is one long blob; and nothing downstream can wait for one particular model. Rendering each dbt node as its own task fixes all four: the Airflow graph mirrors dbt's lineage, a failed model or test retries alone, durations and logs are per model, and each model can be an Airflow asset that other DAGs schedule on (Assets Replace Datasets).

The price is processes. Every task starts dbt, which loads and parses the project before doing any work. On this shared 4-CPU host, one dbt build of Analytical SQL and Data Warehouses's whole project (21 nodes) took 11.4 seconds inside the scheduler container; the same project as 14 Cosmos tasks took 101 seconds, because seven dbt processes started at once and each needed 52-56 seconds of contended CPU. Read the ratio, not the absolute numbers.

Ways to run dbt from Airflow
Approach Airflow sees Cost Run here
BashOperator running dbt build One task One dbt process Yes (Writing DAGs)
Cosmos, ExecutionMode.LOCAL One task per node One process per task Yes
Cosmos, ExecutionMode.WATCHER One task per node One dbt run, watched No
dbt platform job via the dbt Cloud provider One job task Paid dbt platform No

Cosmos's own documentation makes the same trade explicit: per-model tasks give "strong observability and task-level retry control" at the cost of a new dbt process per model, and its watcher mode (stable since 1.15.0) runs dbt once while sensors mark each model's task done. Choose per-model rendering where failures need precise handling; keep a single dbt build for small projects that always succeed or fail as a whole.