Build log
Building a URL shortener in Go that serves 28,000 redirects per second
A production-style URL shortener in Go: two-tier cache, Bloom filter, singleflight and Snowflake IDs, with real benchmarks and the three things to change.
By Winson GR · · 8 min read · Source on GitHub
Most URL shortener write-ups stop at the whiteboard: a box for the API, a box for the database, an arrow labelled “cache”. This one is about the code. The service is a Go API with PostgreSQL and Redis, built around the constraints a real shortener faces, and benchmarked until it stopped scaling. The full source, load tests and Docker setup are on GitHub.
The headline result: 28,093 redirects per second at a p99 of 7.1 ms, with zero errors from 50 to 2,000 concurrent users - and three design decisions that look fine on paper but need revisiting, covered at the end.
The constraints that shape the design
- Reads dominate. A link is created once and opened many times. The design assumes about 100 redirects per new link.
- Traffic is skewed. A small share of links gets most of the clicks, so caching pays off far more than raw database speed.
- A redirect must never wait for anything slow. Click analytics, rate limiting and cache failures must not add latency to the hot path.
Architecture
Every piece exists to keep the redirect path short. The rest of this post walks through them in the order a request meets them.
The read path
A redirect resolves in up to four steps, and each step exists to stop most requests from reaching the next one:
func (s *URLService) Resolve(ctx context.Context, code string, remoteIP, referer, userAgent string) (string, error) {
exists, err := s.bloom.MightExist(ctx, code)
if err != nil {
s.log.Warn("bloom check failed, continuing", zap.Error(err))
} else if !exists {
return "", domain.ErrNotFound
}
u, err := s.l1.Get(ctx, code)
if err != nil {
u, err = s.resolveL2OrDB(ctx, code)
if err != nil {
return "", err
}
_ = s.l1.Set(ctx, u, s.cacheTTL)
}
// expiry checks, then analytics is recorded asynchronously
return u.LongURL, nil
}
Notice two failure-handling choices. If the Bloom filter is unavailable, the request continues instead of failing. And cache writes ignore errors - a cache that can’t be written is a performance problem, not a correctness problem.
Coalescing cache misses with singleflight
When a popular link falls out of cache, hundreds of requests can miss at the same moment and all query the database - a cache stampede. The fix here is golang.org/x/sync/singleflight, which lets only one goroutine per key do the lookup while the others wait for its result:
func (s *URLService) resolveL2OrDB(ctx context.Context, code string) (*domain.URL, error) {
v, err, _ := s.sf.Do(code, func() (any, error) {
u, err := s.l2.Get(ctx, code)
if err == nil {
return u, nil
}
u, err = s.repo.GetByShortCode(ctx, code)
if err != nil {
return nil, err
}
_ = s.l2.Set(ctx, u, s.cacheTTL)
return u, nil
})
if err != nil {
return nil, err
}
return v.(*domain.URL), nil
}
A distributed lock would also prevent the stampede, but costs two Redis round trips per miss. singleflight costs nothing on the network. Its limit is that it only coalesces inside one process: with ten instances, a cold key can still produce ten database queries, which is usually acceptable.
Rejecting codes that don’t exist with a Bloom filter
Without protection, a request for a code that was never created misses L1, misses Redis and queries PostgreSQL. A scraper enumerating codes would turn every guess into a database query. A Bloom filter answers “definitely not here” without touching the database.
The filter is a 10-million-bit Redis bitmap with seven bit positions per code, derived from one SHA-256 hash using Kirsch-Mitzenmacher double hashing. All seven GETBIT calls go out in one pipeline:
func offsets(code string) [numHashes]uint64 {
h := sha256.Sum256([]byte(code))
h1 := binary.BigEndian.Uint64(h[0:8])
h2 := binary.BigEndian.Uint64(h[8:16])
var out [numHashes]uint64
for i := uint64(0); i < numHashes; i++ {
out[i] = (h1 + i*h2) % bitSize
}
return out
}
Using plain SETBIT/GETBIT instead of the RedisBloom module means it runs on any stock Redis. A false positive is cheap: it costs one Redis read and one database read, never a wrong redirect.
The write path
POST /api/v1/shorten
validate scheme http/https, length ≤ 2048
SSRF check resolve DNS, reject private and link-local IPs
normalise lowercase host, sort query params, strip fragment
dedup SHA-256 of the normalised URL → return the existing code if found
generate Snowflake ID → base62 short code
persist INSERT INTO urls
write-through Redis SET + LRU set + Bloom SETBIT ×7
Two decisions are worth calling out.
Normalising before deduplicating. HTTPS://Example.com/a?b=2&a=1#top and https://example.com/a?a=1&b=2 are the same destination, so they get the same short code instead of two rows.
Write-through, not write-behind. The cache is populated before the 201 Created response returns. With write-behind, someone who shares a link the moment it’s created could open it and get a 404. Write-through costs about half a millisecond to close that window.
Short codes come from a Snowflake generator: a millisecond timestamp, a 10-bit worker ID and a 12-bit sequence, so any number of instances can mint IDs without coordinating. Because the IDs only increase, inserts into the primary-key index always land at its right edge instead of splitting pages across the tree.
Keeping analytics off the hot path
Every redirect records a click event, but the redirect never waits for it. Events go into a buffered channel with room for 10,000, and a single goroutine drains it, inserting in batches of 100 or once a second, whichever comes first:
func (r *Recorder) Record(event domain.ClickEvent) {
select {
case r.ch <- event:
default:
r.log.Warn("analytics buffer full, dropping event")
}
}
The default branch is the important line. When the database is slow and the buffer fills, the service drops analytics rather than slowing redirects - an explicit choice that the redirect matters more than the count. Click events go into a table partitioned by time, so old data is removed by dropping a partition instead of running a huge DELETE. IP addresses are stored only as an HMAC with a server-side secret.
Rate limiting runs as a Lua script inside Redis (ZREMRANGEBYSCORE, ZCARD, ZADD in one atomic call) and fails open: if Redis is down, requests are allowed rather than blocked.
Benchmarks
Measured on an Apple M-series laptop with Docker Desktop, three containers with hard CPU and memory limits, rate limiter disabled. Redirects for a cached link:
| Concurrent users | Requests/s | p50 | p95 | p99 | Errors |
|---|---|---|---|---|---|
| 50 | 24,079 | 1.9 ms | 3.6 ms | 4.8 ms | 0 |
| 100 | 28,093 | 3.4 ms | 5.4 ms | 7.1 ms | 0 |
| 200 | 21,682 | 9.1 ms | 13.2 ms | 16.2 ms | 0 |
| 500 | 22,445 | 21.2 ms | 31.4 ms | 41.4 ms | 0 |
| 1,000 | 19,124 | 47.0 ms | 91.9 ms | 138.6 ms | 0 |
| 2,000 | 22,574 | 81.7 ms | 143.0 ms | 221.9 ms | 0 |
Throughput plateaus around 20,000-28,000 requests per second, and past that point latency grows with concurrency instead of throughput. That is exactly what Little’s law predicts for a saturated system: 2,000 users ÷ 22,574 requests/s ≈ 89 ms per request, close to the measured 82 ms median. The extra requests aren’t failing; they’re queueing.
Writes are a different story. Each shorten is a real INSERT plus three cache writes, and tops out at 3,499 requests/s with a p99 of 27.5 ms - the PostgreSQL round trip dominates. That’s still far more than a shortener needs: at 100 reads per write, 28,000 redirects/s implies only 280 new links/s.
What to change before production
Taking the numbers seriously turns up three issues worth fixing before this runs in production.
1. Every redirect pays a Redis round trip, even cache hits. Resolve checks the Bloom filter - seven GETBITs in Redis - before the in-process LRU. So the ~200 ns L1 hit path is never reached without a network call first. Since the Bloom filter only exists to protect the database from codes that don’t exist, it belongs after the L1 check: look in L1 first, and only consult the filter on a miss. Hot links would then be served entirely from memory.
2. The Bloom filter fills up faster than expected. With 10 million bits and 7 hash functions, the false-positive rate is (1 − e^(−kn/m))^k:
| Codes stored | False-positive rate |
|---|---|
| 1 million | ≈ 0.8% |
| 2 million | ≈ 14% |
| 5 million | ≈ 80% |
At 2 million codes, one in seven requests for a non-existent code slips through to the database. Holding 1% at 2 million codes needs about 19 million bits (2.4 MB) - still tiny. The filter needs to be sized for the expected number of links, and rebuilt larger as it grows.
3. Seven-character codes from a 63-bit ID eventually repeat. Encode writes the last seven base62 digits of the Snowflake ID, which keeps only the ID modulo 62⁷ (about 3.5 trillion). Snowflake IDs are far larger than that, so different IDs can share a code. With the same worker and sequence number, two IDs collide when their timestamps differ by a multiple of 31⁷ milliseconds - about 318 days - and some sequence combinations under load collide sooner. The unique index on short_code turns a collision into a failed insert rather than a wrong redirect, but it is still a failure. The fix is either to encode the whole ID (11 base62 characters), or to keep 7 characters, generate them independently of the ID, and retry on a unique-constraint violation.
One more trade-off, deliberate rather than a bug: redirects return 301 Moved Permanently. Browsers cache a 301, which removes load from the server but means repeat clicks from the same browser never reach the analytics pipeline. A shortener that sells click analytics would use 302.
Try the design yourself
The core lesson of this build - a database saturates long before the API does, and a cache in front of it changes everything - is exactly what the simulator lets you feel. Design a URL shortener and watch it break under traffic, then take the URL shortener challenge: grow it from 1,000 to 20,000 redirects per second, survive a cache restart and a server failure, and compare your design with the reference.
The full source, with docker compose up to run everything and the load-test scripts used above, is at github.com/winsongr/url-shortener.