System design · Cache-aside, write-through, stampedes

Caching strategies explained

Almost every system design answer adds a cache. Saying "put Redis in front of it" is easy. The interview questions start after that: how does data get into the cache, how does it get out, what happens when it's wrong, and what happens when it's empty. This page covers the patterns, the trade-offs, and the numbers behind them.

Updated · 6 min read

Why cache

A cache keeps a copy of data somewhere faster or closer than its source. It does two jobs. It cuts latency - a read from memory takes well under a millisecond, while a database query often takes several. And it cuts load - every read the cache answers is a read the database never sees.

The price is two copies of the data that can disagree.

  • Client: the browser or app caches responses, controlled by HTTP headers such as Cache-Control. You can't purge it.
  • CDN: edge servers cache static files and cacheable API responses close to users.
  • Application: an in-process cache inside each server. No network hop, but every server holds its own copy, so they drift apart.
  • Distributed cache: Redis, Valkey or Memcached, shared by all app servers. The usual answer in interviews.
  • Database buffer pool: the database keeps hot pages in memory.

Read strategies

There are two common ways to fill a cache on reads.

Cache-aside, also called lazy loading, puts the application in charge. It checks the cache first. On a miss, it reads the database, writes the result to the cache, and returns it. Only requested data gets cached, and if the cache goes down, reads still work - just slower.

def get_user(user_id):
    key = f"user:{user_id}"
    cached = redis.get(key)
    if cached is not None:
        return deserialize(cached)
    user = db.query("SELECT * FROM users WHERE id = %s", user_id)
    redis.set(key, serialize(user), ex=300)
    return user

def update_user(user_id, fields):
    db.update("users", user_id, fields)
    redis.delete(f"user:{user_id}")

Read-through moves the same logic into the cache layer. The application only talks to the cache, and the cache loads from the database on a miss. In both patterns, the first read of every key is a miss, so a cold cache sends a burst of traffic to the database. Note that the update deletes the key instead of writing the new value. Deleting is safer: two concurrent updates that each write to the cache can finish in the wrong order and leave the old value there.

Write strategies

Writes decide how stale the cache can get and what you lose if something fails.

StrategyHow it worksConsistencyWrite latencyMain risk
Write-throughWrite to cache and database together; return when both succeedStrong between cache and databaseHigher - two writes on the pathCaches data that is never read
Write-behind (write-back)Write to cache, return, flush to database later in batchesEventual - database lags the cacheLowestData loss if the cache fails before flushing
Write-aroundWrite to database only; cache fills on the next readStale until expiry or invalidationSame as no cacheRecently written data always misses first
Cache-aside with deleteWrite to database, then delete the cache keyShort stale windowDatabase write plus one deleteA race can re-cache an old value

Write-behind suits write-heavy data where losing a few seconds of updates is acceptable, such as view counters. Write-around suits data rarely read back, such as logs. For most interview answers, cache-aside with delete-on-write plus a TTL is the safe default.

Expiry and eviction

Expiry and eviction are different. Expiry removes data because it's old. Eviction removes data because the cache is full.

A TTL (time to live) puts an upper bound on staleness. If a key expires after 60 seconds, no reader sees data more than 60 seconds old, even if an invalidation was missed. Short TTLs mean fresher data and more misses. Long TTLs mean more hits and staler data. Pick the TTL from the business answer to "how wrong can this be, and for how long?"

  • LRU (least recently used) evicts the key that hasn't been read for the longest time.
  • LFU (least frequently used) evicts the key with the fewest reads. It protects steadily popular keys from being pushed out by a one-off scan of many cold keys.
  • Invalidation is the hard part. Between the database write and the cache delete there is a stale-read window. Readers in that window see the old value. Keep a TTL as a backstop in case a delete is lost.

The stampede

A cache stampede, or thundering herd, happens when a popular key expires and many requests miss at the same moment. Each one goes to the database to rebuild the same value. If a key gets 5,000 reads a second and the rebuild takes 200 ms, up to 1,000 requests can hit the database for one key before the first rebuild finishes.

  • Request coalescing: let one request rebuild the value while the others wait for it, using a lock or a single-flight helper. One database query instead of a thousand.
  • Jittered TTLs: add randomness, such as 300 seconds plus or minus 30. Keys written together then don't all expire together.
  • Early refresh: rebuild a hot key in the background shortly before it expires, so readers never see a miss.
  • Serve stale: if the rebuild is slow or fails, return the old value for a little longer rather than an error.

You can watch a stampede happen in the caching simulation at /embed/caching-basics, then try to survive one in the /challenges/the-stampede challenge.

Sizing and hit ratio

Hit ratio is the share of reads the cache answers. The database sees the rest. At a 90% hit ratio, the database sees 10% of reads. That is why small drops matter so much.

Hit ratioShare reaching databaseDatabase reads at 50,000 reads/s
99%1%500/s
95%5%2,500/s
90%10%5,000/s
80%20%10,000/s

Dropping from 90% to 80% looks like a 10-point change. It doubles database load. Dropping from 99% to 95% multiplies it by five. Always talk about the miss ratio when you reason about the database. To size the cache, start from the hot data, not the whole dataset. Access is usually skewed: a common rule of thumb is that about 20% of the data gets about 80% of the reads. With 100 million records at 1 KB each, the full dataset is 100 GB. Caching the hot 20% needs about 20 GB, plus overhead for keys and memory fragmentation.

What interviewers look for

  • Naming the pattern you use - usually cache-aside - and explaining how reads fill it and how writes invalidate it.
  • Stating the staleness you accept and the TTL that bounds it.
  • Doing the hit ratio math and showing what the database must handle on misses and on a cold start.
  • Knowing the stampede and at least two fixes for it.
  • Sizing the cache from the hot set, and saying what happens when a cache node fails.

Frequently asked questions

What is the most common caching strategy?

+

Cache-aside. The application reads from the cache, falls back to the database on a miss, and stores the result with a TTL. On writes it updates the database and deletes the cache key.

What is the difference between write-through and write-behind?

+

Write-through writes to the cache and the database before returning, so they stay in sync but every write pays for both. Write-behind writes to the cache and returns, then flushes to the database later. It's faster, but updates are lost if the cache fails before the flush.

Should I update or delete the cache on a write?

+

Usually delete. Two concurrent writers that update the cache can finish in the wrong order and leave an old value cached. Deleting means the next read loads the current value from the database.

How do you prevent a cache stampede?

+

Make sure only one request rebuilds a missing key while others wait, add random jitter to TTLs so keys don't expire together, and refresh hot keys before they expire.

How big should a cache be?

+

Big enough for the hot data. If about 20% of records get most of the reads, size for that 20% plus overhead, then check the hit ratio. Each point of hit ratio you lose adds directly to database load.

Now break one yourself.

The first challenge takes about two minutes. No signup.