Interviews roadmap

Caching, load balancing, and horizontal scaling, briefly

Log in to save this

Saving keeps this in your list across devices. It's a free account — no card.

These three come up in nearly every system design conversation because they solve the same underlying problem — a single machine or database eventually can't keep up — in three different, complementary ways.

Caching

Storing a copy of expensive-to-compute or frequently-read data somewhere faster to access (in memory, closer to the user) so you don't redo the work every time. The cost isn't performance — it's correctness risk: a cache can serve stale data, and cache invalidation (knowing when to throw a cached value away) is a genuinely hard problem, not an afterthought.

Load balancing

Spreading incoming requests across multiple servers instead of one, so no single machine is a bottleneck or a single point of failure. The load balancer itself becomes a component worth reasoning about — what happens if it fails, and how does it decide which server gets the next request (round-robin, least-connections, by request hash).

Horizontal scaling

Adding more machines to handle more load, instead of making one machine bigger (vertical scaling). It's usually the more resilient long-term approach, but it's not free — it typically requires your application to be stateless (any server can handle any request) or requires a real strategy for sharing state across machines (a shared database, a distributed cache).

How they stack

A typical answer under load uses all three together: load balancer spreads requests across app servers, app servers check a cache before hitting the database, and both the app layer and (eventually) the database scale horizontally. Naming this stack, and why each layer is there, is usually enough depth for a first pass at a design.

Resources

Curated resources for this node are on the way. Use what you already know how to search for, and check back soon.