Data architecture is the set of design decisions that are expensive to change later: which storage layer holds the history, batch or streaming ingestion, one central warehouse or domain-owned data products, which cloud. A useful test is to sort decisions into reversible ones, such as trying a new dashboard tool, which you can make quickly, and near-irreversible ones, such as choosing a table format for ten years of history, which deserve a written comparison and a prototype. Lambda and Kappa Architectures and Warehouses to Fabric survey the main architectural patterns.
Orchestration runs the platform's jobs in the right order, at the right time, and handles what goes wrong. An orchestrator models a pipeline as a DAG, a directed acyclic graph of tasks: each task starts only when the tasks it depends on have succeeded, independent tasks run in parallel, and failed tasks are retried and reported.

Cron entries spaced "far enough apart" let a slow extract feed half-loaded data to the next job; an orchestrator replaces guessed gaps with explicit dependencies, retries, backfills and a UI that shows which run failed. Orchestration and Pipelines builds this DAG for real in Apache Airflow 3 129 and compares it with Dagster 177,056 and Prefect 74,615 .