Ingestion moves data from source systems into storage you control. It is where most pipelines break, because it is the stage that touches systems you do not own. Three choices shape every ingestion job:
Batch or streaming. Copy a bounded chunk on a schedule (every night, every hour), or process each record continuously as it arrives. Batch and Streaming Processing compares them.
Push or pull. The source sends data to you (a webhook, an app publishing events to Kafka 129 ), or you fetch it (querying a database, calling an API, listing a bucket for new files).
Full or incremental. Copy everything each time, which is simple and fine for small tables, or copy only what changed since the last run, using an updated_at column, an ever-increasing ID or the database's own change log.
| Pattern | Example at BookNest | Taught in |
|---|---|---|
| Scheduled file pull | Publishers' XML catalog feeds | XML and Its Toolchain and Orchestration and Pipelines |
| Incremental query | New orders since the last run | Pipeline Foundations |
| Change data capture | Every insert and update in the orders table | CDC with Debezium 3.x |
| Event streaming | App events published to Kafka | Apache Kafka and Managed Cloud Kafka |
| Managed connectors | SaaS data loaded by an ingestion tool | Ingestion Tools |
Change data capture (CDC) deserves an early mention. Instead of querying a table again and again, a CDC tool such as Debezium 317,608 reads the database's write-ahead log, the record of every change the database already keeps for its own recovery, and turns each insert, update and delete into an event. The source database barely notices, and deletes are captured too, which incremental queries miss.
Whatever the pattern, expect two kinds of trouble. Networks fail and jobs retry, so the same records will sometimes arrive twice; design the load to tolerate duplicates. Records also arrive late and out of order, so decide early whether "yesterday's orders" means orders placed yesterday or orders received yesterday.