Foundations roadmap

Cache Stampedes and Thundering Herds

Log in to save this

Saving keeps this in your list across devices. It's a free account — no card.

Deep in the Prod Data Center, the database hums along at 10% load behind a cache. Then, once a minute, it spikes to 100% and requests time out, and a moment later everything is calm again. Nothing is broken in the code you wrote last week; the cache is doing exactly what you told it to. This lesson is about that moment when a crowd hits your database all at once, and the small changes that prevent it.

What a stampede is

Take a homepage that shows "top products," cached for 60 seconds:

async function getTopProducts() {
  const hit = await cache.get('top-products');
  if (hit) return JSON.parse(hit);
  const rows = await db.query(TOP_PRODUCTS_SQL); // takes 800ms
  await cache.set('top-products', JSON.stringify(rows), { EX: 60 });
  return rows;
}

With 2,000 requests a second, the cache answers nearly all of them. But when the key expires, every request in the next 800ms finds nothing and runs the same slow query. That is 1,600 copies of an 800ms query at once, and each one slows the others down, so it lasts even longer. This is a cache stampede: the cache protected the database so well that the database can no longer survive without it.

Fix 1: request coalescing (single-flight)

Only one request needs to rebuild the value; everyone else can wait for that same result. Keep the in-progress promise in a map:

const inFlight = new Map();

async function getCached(key, load, ttlSeconds) {
  const hit = await cache.get(key);
  if (hit) return JSON.parse(hit);
  if (inFlight.has(key)) return inFlight.get(key);

  const promise = load()
    .then(async (value) => {
      await cache.set(key, JSON.stringify(value), { EX: ttlSeconds });
      return value;
    })
    .finally(() => inFlight.delete(key));
  inFlight.set(key, promise);
  return promise;
}

The finally matters: if the load fails, the entry is removed and the next request tries again, instead of every request getting the same old error. This map lives in one process, so with ten servers you get up to ten queries instead of thousands, which is usually plenty. If you need exactly one, use a short lock in Redis (SET key value NX EX 10) so only the winner rebuilds.

Fix 2: jittered TTLs

A quieter version of the problem: at startup you warm the cache with 5,000 product pages, all with a 300-second TTL. Five minutes later they all expire together, and again five minutes after that. Add randomness so expiries spread out:

const ttl = 300 + Math.floor(Math.random() * 60); // 300–359 seconds

The same trick helps with scheduled jobs: if every client refreshes "at the top of the minute," spread them out.

Fix 3: serve stale while refreshing

Often a value that is a few seconds old is perfectly fine. Store the data with a "fresh until" time and keep it in the cache longer than that. When a request finds a value past its fresh time, it returns the stale value right away and starts one background refresh (with single-flight, so it's only one). Users always get a fast answer, and the database sees one query per refresh. This is the same idea as the stale-while-revalidate option in HTTP caching headers.

Raising the TTL alone is not a fix. It makes stampedes rarer, but each one is just as bad, and your data is staler the rest of the time.

The same herd in retries

Stampedes aren't only about caches. Say a payment service goes down for a minute, and 5,000 clients retry every second. When it comes back, all 5,000 retries land in the same instant and knock it over again. That is a thundering herd. The fixes follow the same idea:

  • Exponential backoff: wait 1s, then 2s, 4s, 8s, with a cap, so pressure drops over time.
  • Jitter: randomize each wait, such as delay * (0.5 + Math.random()), so clients don't retry in lockstep.
  • A retry limit, so clients eventually give up and report the error.

Spotting it in production

Stampedes have a signature on your dashboards: database load or latency spikes at a regular rhythm that matches a TTL, cache hit rate dips at the same moments, and slow queries in the logs are the same query hundreds of times. When a recovering service falls over again right after coming back, suspect a retry herd. Both point to the same lesson: when many callers can make the same expensive request at the same moment, make sure only one does, or spread them out.

Resources

Curated resources for this node are on the way. Use what you already know how to search for, and check back soon.