Time to move some data. This section takes BookNest's six books through four steps that every analytical pipeline has in some form: a source system hands over a CSV export, a conversion step stores it as typed, compressed Parquet 129 , a SQL query answers a business question, and a chart serves the answer to a person. The data is tiny so that every byte stays visible; the shape is the same at six billion rows. Work in one folder, ~/booknest, with the virtual environment active, and copy the six-book catalog books.json into it (demos/ch01/pipeline/books.json in the book's code).
