Work down this list in order; the early items are cheaper and fix more jobs than memory settings do.
| Check | Where to look | Typical fix |
|---|---|---|
| Reads only needed columns and rows | explain(): ReadSchema, PushedFilters | Select early; partition by date |
| No accidental shuffle or sort-merge join | explain(), SQL tab | Broadcast dimensions; bucket hot joins |
| Balanced tasks | Stages tab: max vs median | AQE skew join; salting |
| No Python UDF on the hot path | Plan: BatchEvalPython | Built-ins, pandas 16,086 or Arrow 129 UDFs |
| Sensible file sizes | File counts, scan time | Coalesce before write; compaction |
| Shuffle partitions | Reduce task count, spill | AQE on; size for the largest shuffle |
| Memory pressure | Spill, GC time, exit code 52 | No collect; more partitions; memory last |
| Repeated work | Jobs per action | Cache reused, expensive inputs |
Make one change per run and measure before and after on the same input; on a shared host, repeat and compare ratios.