System design · Media uploads, feeds, hot accounts

How to design Instagram

Instagram is two systems in one. The first moves heavy files: photos uploaded once, resized, and served billions of times from edge caches. The second is a social feed: who follows whom, and what shows up when you open the app. Most candidates spend all their time on the feed. Interviewers also want to see you keep gigabytes of image bytes away from your application servers.

Updated · 6 min read

Requirements

Pin down the core before drawing boxes. Instagram has stories, reels, direct messages and search - say you're leaving them out.

  • Functional: upload a photo with a caption; follow and unfollow users; view a home feed of recent posts from people you follow; like and comment on posts.
  • Functional, optional: a profile grid of a user's posts; like and comment counts on every post.
  • Non-functional: photos must load fast anywhere in the world; the feed should render in a few hundred milliseconds.
  • Non-functional: an uploaded photo must never be lost. A new post may take a few seconds to reach followers' feeds, and a like count may lag slightly - eventual consistency is fine for both.

Capacity estimates

These are assumptions, stated out loud. The point is the order of magnitude, which decides the architecture.

QuantityAssumptionResult
Uploads500M daily users, 1 in 10 posts a photo a day50M a day ≈ 580/s average, ≈ 1,700/s at a 3× peak
Photo storage≈ 2 MB original + ≈ 500 KB of resized variants≈ 125 TB a day ≈ 45 PB a year, before replicas
Feed reads20 feed loads per user per day10B a day ≈ 115,000/s average
Image requestsabout 10 images per feed load≈ 1.2M images/s, almost all from the CDN
Likes10 likes per user per day5B a day ≈ 58,000/s average

Three conclusions. Image bytes dwarf everything else, so they belong in object storage behind a CDN, never on application servers. Feed reads outnumber uploads by about 200 to 1, so precomputing feeds pays off. And likes are a high-volume write stream that can't update one database row per tap.

API design

Uploads happen in two steps so the photo bytes go straight to object storage. The API only handles small JSON requests.

POST /api/uploads            { "content_type": "image/jpeg", "size": 2048000 }
→ 200 { "upload_id": "up_81", "url": "https://media-bucket.s3.amazonaws.com/...signed" }

PUT  {signed url}             <photo bytes>   (client → S3 directly)

POST /api/posts               { "upload_id": "up_81", "caption": "..." }   → 201 { "id": "1850..." }
POST /api/follows             { "followee_id": 42 }                      → 204
POST /api/posts/{id}/likes                                               → 204
GET  /api/feed?cursor=…        → 200 { "posts": [...], "next_cursor": "…" }

The pre-signed URL is short-lived and scoped to one object key. The client uploads to S3 itself, so a slow mobile connection ties up S3, not your servers.

Data model

Metadata is small and structured; image bytes are large and immutable. Keep them in separate stores. Metadata fits a key-value store such as DynamoDB, or PostgreSQL sharded by user.

posts
  post_id        bigint     time-ordered, e.g. Snowflake
  author_id      bigint     partition key for the profile grid
  media_key      text       S3 object key prefix
  caption        text
  status         text       pending | ready
  created_at     timestamp

follows
  follower_id    bigint     who I follow
  followee_id    bigint     plus a reverse index: who follows me

post_counters
  post_id        bigint
  like_count     bigint
  comment_count  bigint

likes
  post_id        bigint     partition key
  user_id        bigint     sort key, so a second like is a no-op

Store the S3 key, not a full URL. The CDN domain and the variant (thumbnail, feed size, full size) are added when the response is built, so you can change either without rewriting rows.

High-level design

Separate the media path from the metadata path. They scale for different reasons.

  • Upload: the API service issues a pre-signed S3 URL and the client uploads the original directly. The post row is created with status pending.
  • Processing: an S3 event triggers resizing workers - Lambda for bursty load, or EC2 workers reading from SQS for steady volume. They write a few fixed sizes back to S3, strip metadata such as GPS location, and mark the post ready.
  • Delivery: CloudFront sits in front of the S3 bucket. Variants are immutable, so they can be cached at the edge for a long time and almost never hit S3.
  • Feed: when a post is ready, a fan-out service pushes its ID into followers' feed lists in ElastiCache (Redis). A feed read takes the ID list, hydrates posts from a post cache, and adds image URLs pointing at CloudFront.

Where it breaks

Route photo uploads through your API servers and they break first. Each upload holds a connection for seconds on a slow phone, and resizing is CPU-heavy. A few hundred concurrent uploads exhaust the fleet, and feed reads - the thing users actually wait on - start timing out with them.

The fix is to take the bytes off the request path: pre-signed URLs for upload, an asynchronous queue for resizing, a CDN for delivery. The new risk is the window where a post exists but its images don't. Keep the post in pending status, and only fan it out once the resizing workers mark it ready.

Home feed and the celebrity problem

The feed is the same problem as the Twitter timeline, covered in depth in our Twitter timeline guide. The short version: fan out on write for normal accounts, so a feed read is one cached list. For accounts with millions of followers, skip the fan-out and merge their recent posts in at read time. Otherwise one celebrity post becomes millions of Redis writes and delays everyone else's posts.

Likes and comment counters

  • Never run UPDATE like_count = like_count + 1 on every tap. A popular post turns into one hot row with thousands of writes a second.
  • Write the like itself to the likes table (idempotent on post_id and user_id), and increment a counter in Redis.
  • Flush counters to post_counters in batches every few seconds. For very hot posts, split the counter into several shards and sum them on read.
  • The displayed count can lag by seconds. Nobody notices 10,402 versus 10,407.

Caching layers

Every layer of the read path has its own cache. CloudFront serves images. Redis holds feed ID lists, hot post metadata and counters. The database sees cache misses and writes. Size the post cache for the working set - recent posts get almost all views - and let old posts fall out with an LRU policy.

What interviewers look for

  • Keeping image bytes off application servers with pre-signed URLs, and serving them from a CDN.
  • Asynchronous resizing, with a clear answer for posts whose images aren't ready yet.
  • A feed design that names fan-out on write vs read and handles high-follower accounts.
  • Counters that don't turn a viral post into a hot database row.
  • Estimates that separate media bandwidth from metadata QPS, with numbers.

Frequently asked questions

Why use pre-signed URLs for photo uploads?

+

A pre-signed URL lets the client upload straight to S3 with a short-lived, single-object permission. Your servers never handle the bytes, so slow uploads and large files don't consume application capacity, and S3 scales the upload path for you.

Where are Instagram photos stored?

+

In object storage such as S3, with a few resized variants per photo. Metadata - author, caption, S3 key - lives in a separate database. A CDN such as CloudFront serves the images, so most requests never reach S3.

When should images be resized?

+

Asynchronously, right after upload. An S3 event or queue message triggers workers that produce fixed sizes for thumbnails, feed and full view. Producing them up front keeps reads simple and cacheable, at the cost of a little extra storage.

How does the Instagram feed handle celebrities?

+

Most posts are fanned out on write into followers' cached feeds. Posts from accounts with millions of followers are not fanned out; they're fetched and merged in when a follower loads the feed. This hybrid keeps both writes and reads bounded.

How do you count likes at scale?

+

Record each like idempotently, increment an in-memory counter, and write totals to the database in batches. Shard the counter for very hot posts. The count is eventually consistent, which is acceptable for a number users only glance at.

Now break one yourself.

The first challenge takes about two minutes. No signup.