Foundations roadmap

Retries, Backoff and Scheduled Jobs

Log in to save this

Saving keeps this in your list across devices. It's a free account — no card.

Calls to other services fail, and many of those failures are brief: a network hiccup, a server restarting, a moment of overload. Trying again a little later often works. But retrying the wrong way can turn a small hiccup into an outage, or charge a customer twice. At the Scheduler Tower, the clocks decide not just when to try again, but whether to try at all.

Which failures are worth retrying

A retry only helps if the next attempt could succeed with the same request.

  • Worth retrying: timeouts, dropped connections, 503 Service Unavailable, 502/504 from a gateway, and 429 Too Many Requests (after waiting as long as Retry-After says). These describe a temporary state of the other side.
  • Not worth retrying: 400 Bad Request, 401, 403, 404, 422. These say something is wrong with your request. Sending the same request again will get the same answer every time; fix the input or the code instead.

A 500 is a judgement call: sometimes temporary, often a bug. One or two retries is fine; retrying it forever is not.

Exponential backoff with jitter

Retrying instantly usually hits the same problem. Exponential backoff waits longer after each failure: 1 second, then 2, then 4, then 8. That gives the other service time to recover.

Backoff alone has a flaw. If a service blips and 1,000 clients fail at the same moment, they all retry at exactly 1 second, then exactly 2, arriving in synchronised waves that knock the recovering service down again. Jitter adds randomness to each wait, so the retries spread out instead of arriving together.

async function withRetry(fn, maxAttempts = 5) {
  for (let attempt = 1; ; attempt++) {
    try {
      return await fn();
    } catch (err) {
      if (!isRetryable(err) || attempt === maxAttempts) throw err;
      const base = 1000 * 2 ** (attempt - 1);   // 1s, 2s, 4s, 8s
      const delay = Math.random() * base;       // "full jitter"
      await new Promise((r) => setTimeout(r, delay));
    }
  }
}

Retry limits

Always stop eventually. Without a limit, a dependency that is down for an hour means every request piles up retries, each holding memory and connections, and your service goes down too. Pick a small number of attempts (3 to 5 is common), then give up and report the error, or hand the work to a queue to try later. In a chain of services, only one layer should retry; if every layer retries 3 times, the bottom service sees 3 × 3 × 3 = 27 attempts.

Idempotency keys: retrying without doing it twice

A timeout is ambiguous. Your "charge $50" request may have failed before reaching the payment service, or it may have gone through and only the response got lost. Retrying blindly could charge twice.

An idempotency key fixes this. The client generates a unique ID for the operation and sends it with every attempt. The server remembers keys it has already handled, and if it sees one again, it returns the original result instead of doing the work again.

const key = crypto.randomUUID(); // created once, reused on every retry
await withRetry(() =>
  payments.charge({ amount: 5000, customerId }, { idempotencyKey: key })
);

Many payment APIs support this directly. Creating the key inside the retried function would defeat it, because each attempt would look new.

Scheduled jobs that are safe to run twice

Some work runs on a timer rather than a request: a nightly report, deleting expired sessions, sending reminder emails. These are often defined with cron syntax, where 0 3 * * * means "every day at 03:00".

Scheduled jobs get run twice more often than you would think: two server instances each run the scheduler, a deploy restarts mid-run, or someone triggers it manually to catch up. So design them the same way as queue handlers:

  • Make the effect idempotent. "Delete sessions that expired" is naturally safe to repeat. "Send reminders" is not unless you record who already got one today.
  • Work from state, not from the clock. "Process every invoice not yet marked sent" recovers after a missed night; "process invoices created in the last 24 hours" skips or doubles work if a run is late or repeated.
  • Make sure only one instance runs it, for example with a database lock or by running the scheduler in a single place.

Resources

Curated resources for this node are on the way. Use what you already know how to search for, and check back soon.