Memory vs Disk Partitions

Partitions in Memory Versus Partitions on Disk

Partition means two things in Spark 129 . An in-memory partition is the slice of a DataFrame that one task processes, and their number sets a stage's parallelism. File reads pack files into splits of up to spark.sql.files.maxPartitionBytes (128 MB), so the 20 MB order-lines table, stored as four files, arrives as four partitions; every shuffle produces spark.sql.shuffle.partitions (200), which AQE coalesces. A disk partition is a directory such as month=2026-06, written by partitionBy, that lets queries skip files. The two interact: each writing task creates one file per directory it has rows for, so 200 tasks writing 18 months can leave 3,600 small files (Small Files Problem).