A broadcast join collects the small side to the driver, ships one copy to each executor as a hash table, and streams the large side through it with no shuffle and no sort. The listing disables automatic broadcasting, then requests it with F.broadcast(), which overrides the threshold; a noop write runs a plan without output.
import time
def best_of_3(df):
runs = []
for _ in range(3):
t0 = time.perf_counter()
df.write.format("noop").mode("overwrite").save() # run fully, write nothing
runs.append(time.perf_counter() - t0)
return min(runs)
by_country = lambda c: lines.join(c, "customer_id").groupBy("country").agg(F.sum("qty"))
spark.conf.set("spark.sql.autoBroadcastJoinThreshold", "-1") # no automatic broadcast
smj = best_of_3(by_country(customers))
bhj = best_of_3(by_country(F.broadcast(customers))) # the hint still wins
print(f"sort-merge {smj:.2f}s, broadcast {bhj:.2f}s: {smj / bhj:.1f}x")sort-merge 1.21s, broadcast 0.86s: 1.4x
Five runs on this shared 4-core machine gave ratios from 1.4x to 2.4x as other work came and went. The gain is modest because a local "shuffle" only copies files; on a cluster it also sends 1.38 million lines over the network. Broadcast dimension tables, but keep them small: the copy is collected to the driver and held in every executor, so a large one fails with out-of-memory errors or a timeout (spark.sql.broadcastTimeout, 300 s).