Partition means two things in Spark 129 . An in-memory partition is the slice of a DataFrame that one task processes, and their number sets a stage's parallelism. File reads pack files into splits of up to spark.sql.files.maxPartitionBytes (128 MB), so the 20 MB order-lines table, stored as four files, arrives as four partitions; every shuffle produces spark.sql.shuffle.partitions (200), which AQE coalesces. A disk partition is a directory such as month=2026-06, written by partitionBy, that lets queries skip files. The two interact: each writing task creates one file per directory it has rows for, so 200 tasks writing 18 months can leave 3,600 small files (Small Files Problem).
MENU
Memory vs Disk Partitions
Partitions in Memory Versus Partitions on Disk