Caching Strategy

Caching Strategy for BookNest's Pipelines

The nightly job derives three reports from an enriched sales table. Caching the shared input looks like an obvious win; measure it.

Three nightly reports with and without a cached inputJavaScript
import time
sales = (lines.where("status = 'delivered'")
         .join(F.broadcast(books.select(F.col("id").alias("book_id"), "genre")), "book_id")
         .join(F.broadcast(customers.select("customer_id", "country")), "customer_id")
         .withColumn("revenue", F.col("qty") * F.col("unit_price")))
def nightly(base):                               # three reports from one enriched input
    t0 = time.perf_counter()
    for keys in (["genre"], ["country"], [F.date_format("order_ts", "yyyy-MM")]):
        base.groupBy(*keys).agg(F.sum("revenue")).write.format("noop").mode("overwrite").save()
    return time.perf_counter() - t0
nightly(sales)                                   # warm-up
print(f"no cache   {nightly(sales):5.2f}s")
sales.cache()
print(f"cached     {nightly(sales):5.2f}s  (first report fills the cache)")
print(f"cached     {nightly(sales):5.2f}s  (all three from memory)")
sales.unpersist()
Output
no cache   13.65s
cached     17.49s  (first report fills the cache)
cached      3.76s  (all three from memory)

Filling the cache cost more than it saved for the other two reports (four runs: 8.8 to 13.7 s uncached against 13.3 to 17.5 s cached); once filled, the reports took a quarter to a half of the time. So BookNest caches only inputs that several actions reuse and that are expensive to rebuild, caches after filtering, unpersists after the last consumer, shares results between jobs through Parquet 129 or tables (Lakehouses, Data Quality and Governance), and on Spark 4.2.0 129 keeps Arrow-optimized Python UDFs away from cached data and plain Parquet scans (Pandas UDFs (Vectorized UDFs)).