Design a web crawler

hard~45 min#distributed-systems#hld#storage#system-design

Design a crawler that fetches a large portion of the public web, stores the pages, and keeps them reasonably fresh.

Be specific about how you avoid re-crawling the same URL and how you avoid getting your crawler banned.

Solution — locked

Sign in to unlock this one

The solution opens once you explain the idea in your own words and it passes the grader — which needs an account to record. Signing in is free.

Log in