Caching, load balancing, and horizontal scaling, briefly
These three come up in nearly every system design conversation because they solve the same underlying problem — a single machine or database eventually can't keep up — in three different, complementary ways.
Caching
Storing a copy of expensive-to-compute or frequently-read data somewhere faster to access (in memory, closer to the user) so you don't redo the work every time. The cost isn't performance — it's correctness risk: a cache can serve stale data, and cache invalidation (knowing when to throw a cached value away) is a genuinely hard problem, not an afterthought.
Load balancing
Spreading incoming requests across multiple servers instead of one, so no single machine is a bottleneck or a single point of failure. The load balancer itself becomes a component worth reasoning about — what happens if it fails, and how does it decide which server gets the next request (round-robin, least-connections, by request hash).
Horizontal scaling
Adding more machines to handle more load, instead of making one machine bigger (vertical scaling). It's usually the more resilient long-term approach, but it's not free — it typically requires your application to be stateless (any server can handle any request) or requires a real strategy for sharing state across machines (a shared database, a distributed cache).
How they stack
A typical answer under load uses all three together: load balancer spreads requests across app servers, app servers check a cache before hitting the database, and both the app layer and (eventually) the database scale horizontally. Naming this stack, and why each layer is there, is usually enough depth for a first pass at a design.
Resources
Curated resources for this node are on the way. Use what you already know how to search for, and check back soon.