Loading Files into MinIO

Loading BookNest's Files into MinIO Buckets

The bronze zone keeps each source file exactly as it arrived, organized by prefix.

load_booknest.sh: land the raw files in booknest-rawShell
# Land BookNest's raw files, unchanged, in the booknest-raw bucket (the bronze zone).
SRC="/mnt/d/Books/Data Engineering/demos"
mc mb -p lake/booknest-raw
mc cp -q "$SRC/ch02/booknest-catalog.xml" lake/booknest-raw/catalog/ >/dev/null
for f in books.json customers.jsonl orders.jsonl; do
  mc cp -q "$SRC/ch03/data/$f" lake/booknest-raw/shop/ >/dev/null; done
mc cp -q "$SRC/ch03/data/order_events.jsonl" lake/booknest-raw/events/ >/dev/null
mc cp -q "$SRC/ch03/out/orders.parquet" lake/booknest-raw/parquet/ >/dev/null
mc ls -r lake/booknest-raw | sed -E 's/^\[[^]]*\] *//'
Output
Bucket created successfully `lake/booknest-raw`.
3.8KiB STANDARD catalog/booknest-catalog.xml
36MiB STANDARD events/order_events.jsonl
780KiB STANDARD parquet/orders.parquet
1.9KiB STANDARD shop/books.json
574KiB STANDARD shop/customers.jsonl
22MiB STANDARD shop/orders.jsonl

Any S3-aware engine can now read these files in place; DuckDB 61,228 's httpfs extension (httpfs and Object Storage), given a secret for the MinIO 30,943 endpoint, sums s3://booknest-raw/parquet/orders.parquet to the same 3,303,427.30 as BookNest's Lakehouse (demos/ch08/s3/query_raw.sql). But this is still a lake: nothing stops a job from writing orders-v2.parquet beside the first, and a reader of the parquet/ prefix would silently count both. Open Table Formats shows how table formats close that gap.