Foundations roadmap

Circuit Breakers and Failing Fast

Log in to save this

Saving keeps this in your list across devices. It's a free account — no card.

When a service your backend depends on starts failing, the worst thing your code can do is keep calling it and waiting. Every waiting request holds resources, those run out, and soon your service is failing too, and the services that call you after that. A circuit breaker notices that a dependency is broken and stops calling it for a while, so you fail quickly and keep the rest of your app alive. In the Pub/Sub Relay, Breaker the electrician watches every line and pulls the switch before one fault burns down the whole grid.

How one failure spreads

Picture an online shop whose product page calls a reviews service. The reviews service gets overloaded and starts taking 30 seconds to answer. Each product page request now waits 30 seconds, holding a connection and memory. Traffic keeps arriving, so requests pile up until the shop's server has nothing left, and now even pages that do not show reviews stop loading.

This is a cascading failure: one slow part drags down everything that depends on it, layer by layer. Timeouts limit how long each call waits, but if the dependency is down for ten minutes, you still spend a full timeout on every single request. You need a way to stop calling it at all.

The circuit breaker's three states

A circuit breaker sits in front of calls to a dependency and counts failures. It has three states:

  • Closed (normal): calls go through. The breaker counts failures and timeouts.
  • Open: after too many failures in a short time (say, 5 in a row), the breaker trips. For a cooldown period, every call fails immediately without touching the dependency.
  • Half-open: after the cooldown, the breaker lets a single trial call through. If it succeeds, the breaker closes and traffic resumes. If it fails, it opens again for another cooldown.

The names come from electrical circuits: a closed circuit lets current flow; an open one does not.

class CircuitBreaker {
  constructor(fn, { threshold = 5, cooldownMs = 30000 } = {}) {
    Object.assign(this, { fn, threshold, cooldownMs, failures: 0, openedAt: 0, state: 'closed' });
  }
  async call(...args) {
    if (this.state === 'open') {
      if (Date.now() - this.openedAt < this.cooldownMs) throw new Error('circuit open');
      this.state = 'half-open'; // cooldown over: allow one trial
    }
    try {
      const result = await this.fn(...args);
      this.failures = 0;
      this.state = 'closed';
      return result;
    } catch (err) {
      this.failures += 1;
      if (this.state === 'half-open' || this.failures >= this.threshold) {
        this.state = 'open';
        this.openedAt = Date.now();
      }
      throw err;
    }
  }
}

This sketch is for understanding; in a real project, a library such as opossum handles the details, like letting only one trial call through at a time.

Failing fast with a fallback

"Circuit open" is only useful if your code does something sensible with it. Wrap the call and decide what to show instead:

async function getReviews(productId) {
  try {
    return await reviewsBreaker.call(productId);
  } catch {
    return { reviews: [], unavailable: true }; // fallback
  }
}

Failing fast gives the struggling service room to recover, since you are no longer piling requests on it, and gives your users an answer in milliseconds instead of a 30-second spinner.

Graceful degradation

A fallback is a form of graceful degradation: the app keeps doing its main job with a feature reduced or missing. Good fallbacks depend on the feature:

  • Show the last cached copy of the data, perhaps marked "may be out of date".
  • Hide the section entirely (a product page without reviews is still a product page).
  • Use a sensible default, such as generic recommendations instead of personalised ones.

Decide in advance which features are essential. Checkout cannot degrade into "pretend the payment worked"; that must fail clearly. A recommendations panel can simply disappear.

Health checks

A health check is an endpoint, often GET /health, that a load balancer or hosting platform calls every few seconds to ask "can this instance serve traffic?". If it fails, traffic is routed to other instances or the instance is restarted.

Keep it honest and cheap. Check what this instance truly cannot work without, usually its own database connection. Do not fail the health check just because an optional service like reviews is down: then every instance reports unhealthy at once, and the platform takes your entire app offline over a feature your fallback was already handling.

Resources

Curated resources for this node are on the way. Use what you already know how to search for, and check back soon.