System design · Read-heavy scaling, caching
How to design a URL shortener
A URL shortener (bit.ly, TinyURL) turns a long link into a short code and redirects anyone who opens it. It is the most common warm-up question in a system design interview because it looks trivial and hides every read-heavy scaling lesson: a single write is followed by thousands of reads, and the redirect has to be fast every time.
Updated · 6 min read
Run it: click through the fixes and watch the numbers change
- p99 latency
- 1,669ms
- Errors
- 65.7%
- Cost
- $610/mo
The database's 1 read replica can serve 1,000 reads/s but is receiving 2,970. 66% of requests fail. A cache in front of it would absorb most of those reads.
Embed this simulation in your blog, docs or course
Free to embed. Paste this HTML anywhere that accepts an iframe.
Requirements
Start every design by agreeing what the system must do and how well it must do it. Interviewers expect you to ask, not assume.
- Functional: create a short link for a long URL; redirect a short link to its long URL; optionally allow a custom alias and an expiry date.
- Functional, optional: click analytics per link (count, referrer, country).
- Non-functional: redirects must be fast - aim for a p99 under 50ms at the server - and highly available, because a broken short link breaks every page that embeds it.
- Non-functional: short codes should not be guessable in sequence if links are private, and a code must never point at two different URLs.
Capacity estimates
Use round numbers and say them out loud. The goal is the order of magnitude, which decides the architecture.
| Quantity | Assumption | Result |
|---|---|---|
| New links | 100 million per month | ≈ 40 writes/s on average |
| Redirects | 100 reads per write | ≈ 4,000 reads/s average, ≈ 15,000/s at peak |
| Storage per link | ≈ 500 bytes (code, URL, metadata) | 100M × 12 × 5 years × 500 B ≈ 3 TB |
| Code space | 7 characters of base62 | 62⁷ ≈ 3.5 trillion codes |
Two conclusions fall out. Writes are tiny, so the write path barely matters. Reads are 100 times larger and bursty, so the whole design is about serving redirects cheaply.
API design
Two endpoints are enough. Keep the redirect endpoint at the root so short links stay short.
POST /api/links
{ "long_url": "https://example.com/a/very/long/path", "alias": "launch", "expires_at": "2027-01-01" }
→ 201 Created
{ "code": "launch", "short_url": "https://sho.rt/launch" }
GET /{code}
→ 302 Found
Location: https://example.com/a/very/long/pathData model
The access pattern is a single key lookup - code to URL - so the data model is one table keyed by the code. That shape fits a key-value store such as DynamoDB or Cassandra, or a relational table sharded by code.
links
code varchar(10) primary key
long_url text
owner_id bigint nullable
created_at timestamp
expires_at timestamp nullableKeep click events out of this table. Writing a row on every redirect would turn a read-heavy system into a write-heavy one.
Generating short codes
There are two standard approaches, and interviewers want to hear the trade-off between them.
- Counter + base62: give every link a unique integer ID and encode it in base62 (a-z, A-Z, 0-9). There are no collisions by construction. To avoid a single counter becoming a bottleneck, hand each API server a block of IDs (for example 1,000 at a time) from a coordinator, or use a Snowflake-style generator.
- Hash + truncate: hash the long URL (for example SHA-256) and keep the first 7 characters. The same URL always gets the same code, but two URLs can collide, so you must check and retry on collision.
- Sequential counters produce guessable codes. If links can be private, shuffle the ID with a reversible permutation before encoding, or add random characters.
High-level design
The write path is simple: client → load balancer → API server → ID generator → database. The read path carries almost all the traffic and is where the design is won or lost.
- Client → load balancer → API servers → cache → database. A CDN in front is optional; see the cost note below.
- The API server looks the code up in the cache first; only misses go to the database, and the result is written back to the cache.
- The database is sharded by code, so every lookup touches exactly one shard.
Where it breaks
Send redirect traffic straight from the API servers to the database and the database saturates first. In the simulation above, 3,000 redirects per second against a database that can serve 1,000 means two thirds of requests fail, and the rest wait while the database sits near 100% utilization. Adding API servers does nothing - the bottleneck is behind them.
That is the moment interviewers are waiting for: you find the bottleneck from the numbers, not from memory.
Scaling the redirect path
Redirects are idempotent and the code-to-URL mapping almost never changes, which makes them ideal for caching at every layer.
- Cache the mapping in Redis or Memcached. With a 90% hit ratio, only one redirect in ten reaches the database.
- A CDN can absorb a viral link at the edge, but it bills per request. At 20,000 redirects per second, CloudFront's list price comes to over $50,000 a month, many times the cost of the servers behind it, and a redirect is too small to save meaningful bandwidth. At sustained high traffic, cache in memory and let browsers cache 301s instead.
- Choose the redirect code deliberately: 301 (permanent) lets browsers and CDNs cache the redirect, which removes load but hides repeat clicks; 302 (temporary) sends every click back to you, which keeps analytics accurate but costs more traffic.
- Use a short TTL or explicit invalidation for links that can be edited or expire.
Analytics without slowing redirects
If the product needs click counts, record them asynchronously. The redirect handler publishes a small event to a queue (Kafka, Kinesis or SQS) and returns immediately; a separate consumer aggregates counts into an analytics store. A slow analytics pipeline can then never slow a redirect.
What interviewers look for
- You noticed the read-to-write ratio early and designed for reads.
- You can explain base62 and why 7 characters is enough.
- You put a cache in front of the database and can say what hit ratio you expect and why.
- You know the 301 versus 302 trade-off.
- You kept analytics off the critical path.
Frequently asked questions
What is the hardest part of designing a URL shortener?
+
Serving the read load. Redirects outnumber new links by around 100 to 1, so the challenge is keeping redirect latency low under bursty traffic. Caching the code-to-URL mapping in memory is the key move; a CDN helps with viral bursts but gets expensive per request at sustained high traffic.
Should a URL shortener use SQL or NoSQL?
+
Either works. The access pattern is a single key lookup, so a key-value store like DynamoDB scales horizontally with little effort, and a relational database sharded by code works too. The cache in front of the database matters far more than the database choice.
Should a URL shortener use a 301 or 302 redirect?
+
Use 302 if you need accurate click analytics, because browsers request the short URL every time. Use 301 if you want browsers and CDNs to cache the redirect and remove load, accepting that repeat clicks are no longer counted.
How long should a short code be?
+
Seven base62 characters give about 3.5 trillion codes, which is enough for billions of links per year for decades. Six characters give about 57 billion, which is often enough for a smaller service.
How do you avoid two links getting the same code?
+
Generate codes from a unique counter encoded in base62, which cannot collide. If you hash the URL instead, check whether the code already exists and retry with a different salt on collision.