Foundations roadmap

Rate Limiting Requests

Log in to save this

Saving keeps this in your list across devices. It's a free account — no card.

Any public endpoint will eventually be hit far harder than you planned: by a bot guessing passwords, a scraper, or your own frontend stuck in a loop. A rate limit caps how many requests a caller can make in a period of time, so one caller cannot slow the service for everyone or run up your bill. At the Rate Limit Gate, Bucket hands out tokens, and when yours run out you wait.

Why limit at all

Rate limits solve three separate problems:

  • Abuse. A login endpoint without a limit lets an attacker try thousands of passwords a minute. With a limit of, say, 5 attempts per minute per account, guessing becomes impractically slow.
  • Fairness. If one customer's script sends 1,000 requests a second, every other customer's requests queue behind it. A per-customer limit keeps the service usable for everyone.
  • Cost. If an endpoint calls a paid API (an AI model, an SMS provider), every request costs money. A bug that loops can burn through a month's budget in an afternoon; a limit puts a ceiling on that.

Fixed window: the simplest counter

A fixed window counts requests in blocks of time, like "100 per minute, reset at the start of each minute".

const counts = new Map(); // key -> { windowStart, count }

function allowFixedWindow(key, limit, windowMs) {
  const now = Date.now();
  const windowStart = now - (now % windowMs);
  const entry = counts.get(key);
  if (!entry || entry.windowStart !== windowStart) {
    counts.set(key, { windowStart, count: 1 });
    return true;
  }
  entry.count += 1;
  return entry.count <= limit;
}

It is easy to understand, and that matters. Its weakness is the boundary: a caller can send 100 requests at 12:00:59 and another 100 at 12:01:00, which is 200 requests in two seconds while never breaking the "100 per minute" rule.

Token bucket: bursts with a steady refill

A token bucket gives each caller a bucket that holds up to, say, 10 tokens. Every request spends one token. Tokens drip back in at a steady rate, for example 1 per second, up to the bucket's size. If the bucket is empty, the request is refused.

This allows a short burst (10 requests at once is fine; a page might load several things together) while holding the long-run average to the refill rate. It also avoids the fixed-window boundary spike, because tokens refill smoothly instead of all at once. Many production limiters work this way.

Saying no properly: 429 and Retry-After

When a caller is over the limit, answer with status 429 Too Many Requests, not 500 (which says your server broke) or 403 (which says they are never allowed). Add a Retry-After header saying how many seconds to wait, so well-behaved clients back off instead of hammering you.

if (!limiter.allow(req.user.id)) {
  res.set('Retry-After', '30');
  return res.status(429).json({ error: 'Too many requests, try again in 30 seconds' });
}

On the client side, when you receive a 429, wait at least as long as Retry-After says before trying again.

Who you count: per user, per key, or globally

The key you count by decides who gets blocked. Limit per logged-in user or per API key when you can, because that blocks exactly the caller causing the trouble. For anonymous endpoints like login or signup, per IP address is a common fallback, but remember that many people can share one IP (an office, a university), so keep those limits generous, or limit by the target account as well.

A single global limit ("the whole API accepts 1,000 requests per second") protects your servers, but one noisy caller can use it all up and lock everyone else out. Global limits work as a last line of defence, not a replacement for per-caller ones.

Where the counter lives

The Map above lives in one process's memory. That is fine with one server. Once you run three copies of your API behind a load balancer, each copy has its own counts, so a caller whose requests are spread across them effectively gets three times the limit, and the counts reset whenever a server restarts.

With more than one server, keep the counters in a shared store every server can reach, most often Redis, which can increment a counter and set its expiry in one fast operation. Many hosting platforms and API gateways also offer rate limiting built in, which is often the simplest choice.

Resources

Curated resources for this node are on the way. Use what you already know how to search for, and check back soon.