Ingestion: Getting Data In

Ingestion moves data from source systems into storage you control. It is where most pipelines break, because it is the stage that touches systems you do not own. Three choices shape every ingestion job:

Common ingestion patterns and where this book teaches them
Pattern Example at BookNest Taught in
Scheduled file pull Publishers' XML catalog feeds XML and Its Toolchain and Orchestration and Pipelines
Incremental query New orders since the last run Pipeline Foundations
Change data capture Every insert and update in the orders table CDC with Debezium 3.x
Event streaming App events published to Kafka Apache Kafka and Managed Cloud Kafka
Managed connectors SaaS data loaded by an ingestion tool Ingestion Tools

Change data capture (CDC) deserves an early mention. Instead of querying a table again and again, a CDC tool such as Debezium 317,608 reads the database's write-ahead log, the record of every change the database already keeps for its own recovery, and turns each insert, update and delete into an event. The source database barely notices, and deletes are captured too, which incremental queries miss.

Whatever the pattern, expect two kinds of trouble. Networks fail and jobs retry, so the same records will sometimes arrive twice; design the load to tolerate duplicates. Records also arrive late and out of order, so decide early whether "yesterday's orders" means orders placed yesterday or orders received yesterday.