These ten questions cover the whole chapter, and each turns on a format behavior that surprises experienced developers: a duplicate key that resolves differently by engine, numbers that change on the way through jq 133,477 and JSON Schema, keys that collide, a byte order mark that renames a column, a ZIP code that loses its zero, and schema and type changes that fail loudly in one format and silently in another. Every snippet ran on the book's workstation with DuckDB 1.5.5 61,228 , jq 1.8.1 and Python 3.14 with pandas 2.3.3 16,086 , PyArrow 25 129 , fastavro 1.12.2, protobuf 7.36.2 59,573 and jsonschema 4.26.0. Write down what each one prints and why, and only then check Appendix I, which gives the real output, the reason and the section it comes from. Most of the traps have the same shape: the data was valid, the code raised no error, and a value quietly changed on the way through. Those are the bugs that reach production, because no test fails until someone compares the numbers.
Checking your answers
The script below is how the answers were produced. It runs in the folder that holds the two question files, with the book's virtual environment active (Workstation Setup). Questions 3-10 each print one line starting with the question's number, and questions 1 and 2 print one bare line each, so you can compare your notes line by line.
# Run both question files with the book's virtual environment active; compare with Appendix I.
bash questions-1-2.sh
python questions-3-10.pyQuestions
# 1. A line item with a duplicated qty key, read by DuckDB. Which qty comes back?
# (Section 3.1.3)
duckdb -noheader -list -c "SELECT '{\"qty\": 1, \"qty\": 50}'::JSON->>'qty' AS qty"
# 2. jq reads a total written as 100.10. What does each expression print?
# (Sections 3.1.2 and 3.5.1)
echo '{"total": 100.10}' | jq -c '[.total, .total + 0]'import csv, io, json
import fastavro, pandas as pd, pyarrow as pa
from google.protobuf import wrappers_pb2
from jsonschema import Draft202012Validator
# 3. A computed total came out as NaN. What text does Python write, and would
# JavaScript's JSON.parse accept it? Which argument stops Python writing it?
# (Section 3.1.1)
print("3:", json.dumps({"total": float("nan")}))
# 4. A dict with the integer key 1 and the string key "1" goes to JSON and back.
# What comes back? (Sections 3.1.1 and 3.1.3)
print("4:", json.loads(json.dumps({1: "int key", "1": "str key"})))
# 5. Which of three BookNest prices pass {"multipleOf": 0.01}? (Section 3.3.2)
cents = Draft202012Validator({"multipleOf": 0.01})
print("5:", {price: cents.is_valid(price) for price in (14.99, 4.35, 0.3)})
# 6. Excel saved a file as "CSV UTF-8", read here with encoding="utf-8". What are
# row.get("id") and the column names, and which encoding name fixes it?
# (Section 3.7.3)
text = b"\xef\xbb\xbfid,title\r\n1,Salt and Saffron\r\n".decode("utf-8")
row = next(csv.DictReader(io.StringIO(text)))
print("6:", row.get("id"), list(row))
# 7. pandas reads a customer's ZIP code. What value comes back, and how do you keep
# the code as written? (Section 3.7.2)
print("7:", pd.read_csv(io.StringIO("customer_id,zip\n1,02134\n"))["zip"][0])
# 8. A producer adds the channel "kiosk"; a consumer still has the old enum. What
# happens when it reads the new record, and what could the old enum have declared
# to survive it? (Section 3.9.3)
def order(*symbols):
channel = {"type": "enum", "name": "Channel", "symbols": list(symbols)}
return fastavro.parse_schema({"type": "record", "name": "Order",
"fields": [{"name": "channel", "type": channel}]})
new, old = order("ios", "android", "web", "kiosk"), order("ios", "android", "web")
buf = io.BytesIO()
fastavro.schemaless_writer(buf, new, {"channel": "kiosk"})
buf.seek(0)
try:
print("8:", fastavro.schemaless_reader(buf, new, old))
except Exception as e:
print("8:", type(e).__name__)
# 9. How many bytes does a protobuf Int64Value of 0 take, and what does parsing an
# empty byte string give? (Section 3.10.1)
zero = wrappers_pb2.Int64Value(value=0).SerializeToString()
print("9:", len(zero), wrappers_pb2.Int64Value.FromString(b"").value)
# 10. An Arrow int64 column holding 1 and a null goes to pandas. What values and
# dtype come back, and how do you keep the integers? (Section 3.15.4)
qty = pa.table({"qty": pa.array([1, None], pa.int64())}).to_pandas()["qty"]
print("10:", qty.tolist(), qty.dtype)