The Pipeline as Assets

Modeling BookNest's Pipeline as Assets

The last asset, mart/daily_genre_sales, depends on fct_sales and dim_book and calls publish_day() for its partition. run/final.sh rebuilds everything from scratch: pipeline/setup.sh resets the shop tables to 27 June, up.sh starts l2-dagster (UI on port 32300) with a fresh instance, and run/d3.sh starts booknest_daily three ways: the sensor for 28 June, dagster job launch with the tag dagster/partition=2026-06-29, and a backfill of 30 June.

The upstream lineage of mart/daily_genre_sales in the Dagster UI: ingest assets, dbt models, and the mart
The upstream lineage of mart/daily_genre_sales in the Dagster 177,056 UI: ingest assets, dbt 37,942 models, and the mart

runs.py lists the instance's runs:

Output of 36
3adb617f  booknest_daily        2026-06-28 sensor   SUCCESS 147.3 s
a5b61eb0  booknest_daily        2026-06-29 manual   SUCCESS 157.6 s
2aeab903  booknest_daily        2026-06-30 backfill SUCCESS 101.6 s

Every check passed. But 100-160 seconds is slow. steps.py showed why: the default multiprocess executor starts a new Python process per step, and each one imports Dagster, dagster-dbt and the code location before working. Relaunching 30 June with the in-process executor (--config-json '{"execution": {"config": {"in_process": {}}}}'):

Output of 36
landing__orders_file                 STEP_SUCCESS +  0.2 s    0.7 s
shop_tables                          STEP_SUCCESS +  1.0 s    0.4 s
sales_facts                          STEP_SUCCESS +  1.4 s   12.3 s
fct_sales_fct_sales_matches_source   STEP_SUCCESS + 13.7 s    0.2 s
mart__daily_genre_sales              STEP_SUCCESS + 13.9 s    0.2 s

The run took 14.1 seconds against 101.6, a seventh, on a 4-CPU host shared with two Kafka 129 clusters. Keep process isolation for heavy steps; give chains of quick SQL steps executor_def=dg.in_process_executor.