Files Aren't a Table

Why a Folder of Parquet Files Isn't a Table

If a table is "every file under a prefix", its state changes one PUT at a time. Below, a monthly export's first attempt crashes after 7 of 18 files, the retry writes all 18 under new names, and DuckDB 61,228 counts the folder.

folder_table.py: a reader sees a half-finished job, then a retry's duplicatesPython
"""A folder of Parquet files is not a table: readers see half-finished jobs and duplicates."""
import duckdb
con = duckdb.connect()
con.sql("""SET TimeZone = 'UTC';
  CREATE SECRET (TYPE s3, KEY_ID 'booknest-admin', SECRET 'booknest-secret-2026',
    ENDPOINT 'localhost:31900', URL_STYLE 'path', USE_SSL false, REGION 'us-east-1');
  CREATE VIEW src AS
    FROM '/mnt/d/Books/Data Engineering/demos/ch03/out/orders.parquet'""")
months = [r[0] for r in con.sql(
    "SELECT DISTINCT strftime(order_ts, '%Y-%m') AS m FROM src ORDER BY m").fetchall()]
def write(month, name):
    con.sql(f"""COPY (FROM src WHERE strftime(order_ts, '%Y-%m') = '{month}')
                TO 's3://booknest-demo/folder/orders/{name}.parquet'""")
def reader(label):
    n = con.sql("SELECT count(*) FROM 's3://booknest-demo/folder/orders/*.parquet'").fetchone()
    print(f"{label:<34} {n[0]:>7,} orders")
for m in months[:7]:              # attempt 1 crashes after 7 of 18 files...
    write(m, f"attempt1-{m}")
reader("reader while attempt 1 runs:")
for m in months:                  # ...attempt 2 rewrites everything with new file names
    write(m, f"attempt2-{m}")
reader("reader after the successful retry:")
Output
reader while attempt 1 runs:        38,866 orders
reader after the successful retry: 138,866 orders

Both answers are wrong, and nothing in storage can tell the reader so: a partial write is visible, an abandoned one is permanent, and concurrent writers would lose updates. A table needs an authoritative list of its files that changes in one step, which is what a table format's metadata provides.