Data Engineering Foundations

Every report, dashboard and machine-learning model depends on data that someone moved, cleaned, checked and stored before anyone looked at it. Data engineering is that work: building and running the systems that take data from where it is born, such as an order form, a payment service or a supplier's catalog feed, and deliver it, trustworthy and on time, to the people and programs that use it.

This chapter is the map for the rest of the book. It names the ideas that come up again and again and gives each a page or two and a pointer to the chapter that teaches it properly.

You also meet BookNest, the small online bookshop that runs through all eight chapters. Its six-book catalog becomes XML in XML and Its Toolchain, columnar files in JSON, Columnar and Binary Formats, a sales mart in Analytical SQL and Data Warehouses, a Spark 129 job in Batch Processing with Apache Spark, a stream of order events in Apache Kafka and Managed Cloud Kafka, an orchestrated pipeline in Orchestration and Pipelines and a lakehouse in Lakehouses, Data Quality and Governance. Here you set up the book's workstation (WSL2 6 with Ubuntu 26.04 225 , Docker 514 , Python and DuckDB 61,228 ) and push the catalog through a first, tiny pipeline from CSV to Parquet 129 to SQL to a chart.

What you will learn

Sections