Retries and Caching

Retries, Caching and Result Persistence

Retries are task options, and retry_delay_seconds accepts a list for a back-off schedule. Caching needs a cache key: the default policy hashes the task's inputs, its source code and the flow run ID, so it only helps within one run; INPUTS + TASK_SOURCE drops the run ID so later runs can reuse a result. A cached value must be stored, so caching implies result persistence (off by default, PREFECT_RESULTS_PERSIST_BY_DEFAULT=false):

flows/booknest_flow.py (excerpt): a cached extract and a retried loadPython
@task(cache_policy=INPUTS + TASK_SOURCE, cache_expiration=timedelta(hours=12),
      persist_result=True)                       # same day, same export, same code: reuse
def extract(ds: str, export_mtime: float) -> str:
    path, n = p.extract_orders(ds, EXPORT, f"{DATA}/landing")
    print(f"extracted {n} orders for {ds}")
    return path
@task(retries=2, retry_delay_seconds=[10, 30])   # ride out a database restart
def load(ds: str, path: str) -> dict:
    orders, lines = p.load_orders(ds, path, connect())
    return {"orders": orders, "lines": lines}

export_mtime is in the key on purpose: if the order service rewrites its export, the modification time changes and the cache misses, where a key on ds alone would quietly reuse a stale file. run/p2.sh stopped l2-pg, started 29 June, waited for the first failure and started the database again:

Output of 39
02:12:01 l2-pg stopped
02:12:23 l2-pg started
dangerous-panda  {'ds': '2026-06-29'}  Completed  0:00:20.707008
02:12:17  extract-7a1    Finished in state Cached(type=COMPLETED)
02:12:17  load-aba       Task run failed with exception: OperationalError('could not translate
                         host name "l2-pg" to address: ...') - Retry 1/2 will start 10 ...
02:12:27  load-aba       Finished in state Completed()
02:12:27  flow           {'orders': 192, 'lines': 261}

extract came from the cache of an earlier, failed run of the same day, whose load had used up both retries while the database stayed down. Each result is a small JSON file under PREFECT_LOCAL_STORAGE_PATH, named after its cache key; set a remote result storage block (S3, GCS, Azure 6 ) when workers run on several machines:

Output of 39
{"metadata":{"storage_key":"/opt/prefect/results/7af8fe3ee3a1091407a34280e13ff4ca",
 "expiration":"2026-10-02T14:03:49.138738Z","serializer":{"type":"pickle",...},
 "prefect_version":"3.8.7","storage_block_id":null},
 "result":"gAWVQQAAAAAAAACMPS9vcHQvcHJlZmVjdC9..."}