Caching and Redis

Caching In Process and in Redis

Caching is the one optimization that can buy an order of magnitude without touching the code that produces the answer. It also introduces the two hardest bugs in production software — stale data, and a cache that never evicts — so the real decisions are about eviction and invalidation, not lookups.

An in-process cache is a Map you must never let grow forever. lru-cache 5,917 (Blue Oak 1.0.0, 11.5.2, github.com/isaacs/node-lru-cache (https://github.com/isaacs/node-lru-cache 5,917 ), npm 2,036 i lru-cache) bounds it by entry count, computed size or age, evicting the least recently used entry when a bound is hit. Cache the serialized response, not the object.

An LRU cache in front of the search handlerJavaScript
import { LRUCache } from 'lru-cache';
const cache = new LRUCache({ max: 500, ttl: 30_000 });
function respond(req, res, category) {
  let body = cache.get(category);
  if (body === undefined) cache.set(category, body = search(category));
  res.writeHead(200, { 'content-type': 'application/json' });
  res.end(body);
}

Rerunning the ten-second, fifty-connection test from Load Testing an HTTP Server against this third mode on the same machine gives the last row below. Five categories rotated, so the server reported 158,417 hits and 5 misses.

Measured throughput of one endpoint, 50 connections for 10 seconds
Mode Throughput p50 p99 vs. naive
naive: parse per request 244 req/s 204 ms 245 ms 1x
warm: parse once at startup 8,507 req/s 5 ms 13 ms 35x
warm plus LRU of rendered responses 15,838 req/s 2 ms 7 ms 65x

The cache nearly doubled an already-fixed server by removing the remaining filter and JSON.stringify. Note the order: caching the naive server would have hidden the bug instead of fixing it, and would still have stalled the moment a cold key arrived. Fix first, cache second.

An in-process cache is per process: eight cluster workers (Clustering Across CPU Cores) mean eight copies and eight chances to serve different stale data, though a hit costs tens of nanoseconds and nothing can fail. Redis 2,763 moves the cache out of the process so every worker and container shares one copy that survives restarts, at the price of a round trip and a dependency that can go down. Use ioredis 15,343 (MIT, 6.0.0, github.com/redis/ioredis (https://github.com/redis/ioredis 15,343 ), Node 20 or newer).

The same handler against a shared Redis cacheJavaScript
const redis = new Redis(process.env.REDIS_URL ?? 'redis://127.0.0.1:6379');
async function respond(req, res, category) {
  const key = `search:v3:${category}`;                 // version the key, not the value
  let body = null;
  try { body = await redis.get(key); } catch { /* cache down: fall through */ }
  if (body === null) redis.set(key, body = search(category), 'EX', 30).catch(() => {});
  res.end(body);
}

Three details separate a working Redis cache from an outage. Put a version in the key so a deploy that changes the response shape misses cleanly instead of serving last week's schema; flushing a shared cache to fix that is how you take the database down with it. Always set an expiry. And treat Redis as optional, as above — a cache that can take the site down with it is worse than no cache. Most systems end up with both layers, which also blunts the stampede when a popular key expires and every worker recomputes it at once.