If a table is "every file under a prefix", its state changes one PUT at a time. Below, a monthly export's first attempt crashes after 7 of 18 files, the retry writes all 18 under new names, and DuckDB 61,228 counts the folder.
"""A folder of Parquet files is not a table: readers see half-finished jobs and duplicates."""
import duckdb
con = duckdb.connect()
con.sql("""SET TimeZone = 'UTC';
CREATE SECRET (TYPE s3, KEY_ID 'booknest-admin', SECRET 'booknest-secret-2026',
ENDPOINT 'localhost:31900', URL_STYLE 'path', USE_SSL false, REGION 'us-east-1');
CREATE VIEW src AS
FROM '/mnt/d/Books/Data Engineering/demos/ch03/out/orders.parquet'""")
months = [r[0] for r in con.sql(
"SELECT DISTINCT strftime(order_ts, '%Y-%m') AS m FROM src ORDER BY m").fetchall()]
def write(month, name):
con.sql(f"""COPY (FROM src WHERE strftime(order_ts, '%Y-%m') = '{month}')
TO 's3://booknest-demo/folder/orders/{name}.parquet'""")
def reader(label):
n = con.sql("SELECT count(*) FROM 's3://booknest-demo/folder/orders/*.parquet'").fetchone()
print(f"{label:<34} {n[0]:>7,} orders")
for m in months[:7]: # attempt 1 crashes after 7 of 18 files...
write(m, f"attempt1-{m}")
reader("reader while attempt 1 runs:")
for m in months: # ...attempt 2 rewrites everything with new file names
write(m, f"attempt2-{m}")
reader("reader after the successful retry:")Output
reader while attempt 1 runs: 38,866 orders reader after the successful retry: 138,866 orders
Both answers are wrong, and nothing in storage can tell the reader so: a partial write is visible, an abandoned one is permanent, and concurrent writers would lose updates. A table needs an authoritative list of its files that changes in one step, which is what a table format's metadata provides.