System design guides you can run
How to design the systems engineers and interviewers ask about most - URL shorteners, news feeds, CDNs, rate limiters and AI pipelines. Each guide explains where the design breaks and what fixes it, and the ones marked “Runnable” let you click through the fixes on the live simulator.
Classic system design
- Design an URL ShortenerRunnableRead-heavy scaling, cachingDesign a URL shortener like bit.ly: requirements, capacity estimates, API, data model, short codes, caching, and when a CDN is worth it. Run the design live.
- Design a Chat SystemQueues, write scaling, delivery lagDesign a chat system like WhatsApp: WebSockets, message queues, delivery receipts, group chat fan-out, presence and storage. Then load-test your design.
- Design a Notification SystemQueues, provider rate limits, retriesDesign a notification system: push, email and SMS, a queue in front of workers, provider rate limits, retries, dedup and user preferences. Then load-test it.
- Design a Twitter TimelineRunnableFan-out and the celebrity problemDesign Twitter's home timeline: estimates, fan-out on write vs fan-out on read, the celebrity problem, caching and the hybrid design. Run it live.
- Design an InstagramMedia uploads, feeds, hot accountsDesign Instagram: photo uploads with pre-signed URLs, resizing workers, CDN delivery, the follow graph, feed fan-out and like counters. Then load-test it.
- Design a YouTubeTranscoding pipeline, object storage, CDNDesign YouTube: video upload, a transcoding pipeline with queues and workers, object storage, CDN delivery, adaptive bitrate streaming and view counts.
- Design a Netflix CDNRunnableEdge caching and the launch stampedeDesign video delivery like Netflix: bandwidth estimates, adaptive bitrate, edge caching, origin shield and the launch-day cold-cache stampede. Run it live.
- Design a Rate LimiterToken bucket, distributed stateDesign a distributed rate limiter: token bucket vs sliding window, atomic counters in Redis, race conditions, response headers and failure modes.
- Design an UberGeospatial indexing, matchingDesign a ride-hailing system like Uber: driver location ingestion, geohash, S2 and H3 indexing, nearby search, matching, trip state and surge pricing.
- Design a Web CrawlerPoliteness, dedup, the URL frontierDesign a web crawler: URL frontier, politeness queues, robots.txt, DNS caching, Bloom filters, simhash dedup, recrawl scheduling and crawler traps.
- Design a DropboxChunking, dedup, syncDesign Dropbox or Google Drive: file chunking, block dedup, pre-signed uploads to S3, a metadata service, sync with a change journal, and conflict handling.
- Design a Key-Value StorePartitioning, replication, quorumsDesign a distributed key-value store: consistent hashing, replication, quorum reads and writes, conflict resolution, hinted handoff, Merkle trees and LSM trees.
- Design a Consistent HashingRebalancing with minimal movementConsistent hashing explained: why hash mod N breaks when nodes change, the hash ring, virtual nodes, replication, and alternatives like rendezvous hashing.
- Design a Caching StrategiesCache-aside, write-through, stampedesCache-aside, read-through, write-through, write-behind and write-around explained, plus TTLs, eviction, stampedes and hit ratio math for system design.
AI system design
- Design a RAG PipelineToken budgets, semantic cacheDesign a production RAG system: token and storage estimates, ingestion, vector search, reranking, permissions, and latency and cost per query.
- Design an AI Agent SystemOrchestration, loops, cost per stepDesign a multi-agent AI system: token and latency estimates, orchestrator and workers, tool and model gateways, durable state, loop limits and budgets.