Small Files Problem

File Sizing and the Small Files Problem

Every file costs a listing call, an open and a footer read before any data flows, which Spark 129 models as spark.sql.files.openCostInBytes (4 MB). Thousands of small files, typically from a write after a 200-partition shuffle, a partitionBy on a busy column or a streaming job, make reads crawl.

The same order lines as 2,000 small files and as 4 filesJavaScript
import os, time
def files_and_scan(path):
    n = len([f for f in os.listdir(path) if f.endswith(".parquet")])
    best = min(timed_count(path) for _ in range(3))
    df = spark.read.parquet(path)
    print(f"{path:<16} {n:>5} files  {df.rdd.getNumPartitions():>3} read partitions  "
          f"count {best:5.2f}s")
def timed_count(path):
    t0 = time.perf_counter()
    spark.read.parquet(path).where("qty > 1").count()
    return time.perf_counter() - t0
lines.repartition(2000).write.mode("overwrite").parquet("out/small_files")   # 2,000 tiny files
lines.coalesce(4).write.mode("overwrite").parquet("out/right_sized")          # 4 files
for path in ("out/small_files", "out/right_sized"):
    files_and_scan(path)
Output
out/small_files   2000 files   63 read partitions  count 14.55s
out/right_sized      4 files    4 read partitions  count  0.89s

Spark packed the 2,000 files into 63 read partitions, but each task still opened about 32 files and the scan took 16 times as long. Aim for files of 128 MB to 1 GB: coalesce before writing, use the REBALANCE hint with AQE, and compact tables written by streams (Lakehouses, Data Quality and Governance).