Data is born in source systems: the transactional database behind a web store, the event stream an app emits when someone taps "Buy", a partner's REST API, a nightly CSV export from an accounting package, a supplier's XML catalog feed, the logs of a web server, the readings of a sensor. A data engineer rarely owns these systems, which is the first lesson of the job: the data's shape, timing and quality are decided by someone else, usually a software team with its own priorities.
So before you build anything, interview the source. How much data does it produce, and how fast? Is it a table, a stream or a file? Are updated and deleted values kept or lost? Who tells you when the schema changes? Write the answers down.
BookNest has three kinds of source: its storefront database (customers and orders), the order events its app emits, and the catalog metadata publishers send as XML. To give the later stages something concrete, this chapter's demo generates eight sample order events from the six-book catalog with a seeded Python script (demos/ch01/lifecycle/make_events.py), written as JSON Lines, one event per line:
head -n 3 order-events.jsonl{"order_id": 1001, "book_id": 3, "qty": 2, "unit_price": 24.0, "ts": "2026-09-28T09:22:00Z"}
{"order_id": 1002, "book_id": 6, "qty": 1, "unit_price": 21.3, "ts": "2026-09-28T09:31:00Z"}
{"order_id": 1003, "book_id": 5, "qty": 2, "unit_price": 16.2, "ts": "2026-09-28T09:46:00Z"}Each event records a fact that already happened, with a timestamp in UTC. That makes events easy to append and replay, a property Apache Kafka and Managed Cloud Kafka builds on. JSON Lines itself is covered in JSON Lines.