Every file costs a listing call, an open and a footer read before any data flows, which Spark 129 models as spark.sql.files.openCostInBytes (4 MB). Thousands of small files, typically from a write after a 200-partition shuffle, a partitionBy on a busy column or a streaming job, make reads crawl.
import os, time
def files_and_scan(path):
n = len([f for f in os.listdir(path) if f.endswith(".parquet")])
best = min(timed_count(path) for _ in range(3))
df = spark.read.parquet(path)
print(f"{path:<16} {n:>5} files {df.rdd.getNumPartitions():>3} read partitions "
f"count {best:5.2f}s")
def timed_count(path):
t0 = time.perf_counter()
spark.read.parquet(path).where("qty > 1").count()
return time.perf_counter() - t0
lines.repartition(2000).write.mode("overwrite").parquet("out/small_files") # 2,000 tiny files
lines.coalesce(4).write.mode("overwrite").parquet("out/right_sized") # 4 files
for path in ("out/small_files", "out/right_sized"):
files_and_scan(path)Output
out/small_files 2000 files 63 read partitions count 14.55s out/right_sized 4 files 4 read partitions count 0.89s
Spark packed the 2,000 files into 63 read partitions, but each task still opened about 32 files and the scan took 16 times as long. Aim for files of 128 MB to 1 GB: coalesce before writing, use the REBALANCE hint with AQE, and compact tables written by streams (Lakehouses, Data Quality and Governance).