Transformations Versus Actions

DataFrame methods fall into two groups. Transformations (select, where, withColumn, join, groupBy with agg, orderBy) return a new DataFrame and only extend the plan. Actions (count, collect, first, take, show, toPandas, every write) need a result, so they run jobs. The listing counts each step's jobs.

Counting the jobs each call launchesPython
sc = spark.sparkContext
def step(label, fn):
    sc.setJobGroup(label, label)                  # tag the jobs this step launches
    result = fn()
    print(f"{label:<22} jobs: {len(sc.statusTracker().getJobIdsForGroup(label))}")
    return result
lines = step("read.parquet", lambda: spark.read.parquet("data/order_lines.parquet"))
top = step("where/groupBy/orderBy", lambda: lines.where("status = 'delivered'")
           .groupBy("book_id").agg(F.sum("qty").alias("copies")).orderBy(F.desc("copies")))
step("count()", top.count)
step("first()", top.first)
step("collect()", top.collect)
Output
read.parquet           jobs: 1
where/groupBy/orderBy  jobs: 0
count()                jobs: 3
first()                jobs: 2
collect()              jobs: 4
The count(), first() and collect() jobs in the Spark UI: each action rescans the order lines with 4 tasks
The count(), first() and collect() jobs in the Spark 129 UI: each action rescans the order lines with 4 tasks

The three transformations launched nothing. Reading Parquet 129 ran one small job to fetch the footer's schema. Each action then ran the whole query again: jobs 1, 4 and 6 are the same four-task scan of 1.38 million rows. Under AQE one action becomes several jobs, one per shuffle stage, and "skipped" stages are shuffle output reused within that action. Nothing survives between actions unless you cache it (Caching and Persistence Levels).