Design a retry policy

medium~25 min#distributed-systems#lld#reliability

A service calls a flaky downstream dependency. Design the retry behaviour.

Cover which failures to retry, the backoff, and how to stop your retries from being the reason the downstream never recovers.

Solution

What to retry. Only what might succeed on a second attempt. Timeouts, connection failures, 429, and 5xx are retryable. 4xx other than 429 are not — a 400 will be a 400 forever, and retrying it just multiplies load. Retrying a non-idempotent operation risks duplicating it, so either the operation carries an idempotency key or it must not be retried at all. This distinction is the first thing to state.

Backoff must be exponential and jittered. Exponential alone still synchronises: if a downstream blips and a thousand callers all back off 1s, 2s, 4s, they retry in lockstep and arrive as three thundering herds, each capable of knocking the recovering service over again. Full jitter — sleep a uniform random amount in [0, base * 2^attempt] — spreads them out. The jitter is not a refinement; it is the part that makes backoff work.

Cap everything. A maximum delay, a maximum attempt count, and — most importantly — an overall deadline for the logical operation. Without a deadline, retries inside retries multiply: three layers each retrying three times is 27 requests and a latency that has nothing to do with any configured timeout. Passing a remaining-time budget down the call chain is the fix, and mentioning that layered retries multiply is a strong signal.

Circuit breaker. Retries are the wrong tool for a downstream that is down rather than flaky. Track the recent failure rate; once it crosses a threshold, open the circuit and fail immediately without calling at all. After a cooldown, allow a trial request (half-open) and close on success. This is what stops your retries from being the load that prevents recovery.

Observability. Count retries and circuit state as first-class metrics. A system where retries have silently become 40% of all traffic looks healthy on a success-rate dashboard right up until it does not.

The follow-up they will ask

Three services each retry three times. What is the actual request amplification, and how do you bound it?