Stop the Scraper Flood – mid-level system design challenge
A scraper is about to hit your public API with 10× normal traffic. Keep real users' requests fast while the flood lasts, and be healthy again the moment it stops.
Requirements
Normal traffic: 2,000 req/s
Scraper flood: 20,000 req/s for a while
p99 latency: ≤ 200 ms, even during the flood
Errors: ≤ 0.5% normally; rejected flood traffic is expected
Goal
Hold p99 latency under 200ms, error rate under 0.5% through every stage of the test plan, then compare your design with a reference design.
The test plan
Normal: 2,000 requests per second.
Scraper flood: 20,000 requests per second.
Flood over: 2,000 requests per second.
Hint
You can't out-scale a flood you don't control, and you shouldn't try. Put an API Gateway in front with a rate limit a little above normal traffic: excess requests get an instant HTTP 429 instead of queueing behind real users. Leave the servers behind the limit some headroom.
Solve it on the live simulator, free and in your browser.
More system design challenges
Hello, Load Balancer - One API server is drowning under 1,500 RPS. Get error rate to zero without touching the traffic.
Cache Money - Your database is at 95% utilization serving reads it has answered a thousand times. Cool it down below 50%.
Design a URL Shortener - Redirect traffic grows from 1,000 to 20,000 requests per second. Keep every redirect fast through a cache restart and a server failure at peak, within budget.
The Stampede - Your cache will be flushed mid-run. Survive the stampede - keep errors under 5% while it refills.
Queue It Up - Write traffic spikes 5× partway through. Absorb the burst without dropping a single write.
Design a Chat System - Messages and history reads grow to 10,000 a second, then New Year's midnight triples it. Keep sends fast and every message delivered within 10 seconds, within budget.
Design a Notification System - Your app sends push notifications through a provider that accepts 2,000 a second. Breaking news is about to multiply events 20×. Deliver every notification within a minute without the provider rejecting any.
Design a News Feed - Timeline reads and new posts grow to 30,000 requests a second at peak. Keep timelines fast, get every post into followers' feeds and search within 10 seconds, within budget.
The Retry Storm - Your database is about to slow down by 400 ms. Your clients retry failures up to three times. Keep the API standing.
Black Friday - Traffic is 12,000 RPS and the CFO capped infra at $3,000/month. Keep p99 under 300ms and errors under 1%.
Chaos Monkey - A monkey will kill your busiest API server mid-run. Design for redundancy so users never notice.
The CTO Budget Cut - This system works - and burns about $15,000/month doing it. Hit the same SLOs for under $1,500/month.