System design · Chunking, dedup, sync

How to design Dropbox or Google Drive

File sync looks like a file server with a nice client. The hard parts are elsewhere: not re-sending what the server already has, telling every device about a change within seconds, and handling two people editing the same file - one of them offline on a plane.

Updated · 7 min read

Requirements

Scope it first. Interviewers want to see you pick the core and say what you're leaving out.

  • Functional: upload, download, move and delete files; changes sync to all of a user's devices; old versions can be restored.
  • Functional: share a file or folder with other users as viewer or editor; edits made offline sync when the device reconnects.
  • Non-functional: a saved file must never be lost or silently corrupted; a change on one device should appear on others within seconds.
  • Non-functional: editing a few bytes of a large file should not re-upload the whole file. Out of scope: real-time co-editing, search, previews.

Capacity estimates

These are assumptions, stated out loud. The point is the order of magnitude, which decides the architecture.

QuantityAssumptionResult
File changes100M daily users × 10 saved changes each1B a day ≈ 11,600/s average, ≈ 35,000/s at a 3× peak
Upload bandwidthabout 1 MB of new data per change≈ 1 PB a day ≈ 12 GB/s average, before dedup and delta sync
Stored data500M accounts × 2 GB used on average≈ 1 EB logical, before replicas and dedup
Sync connectionshalf of daily users have a device online at peak≈ 50M open long-poll or notification connections

Three conclusions. File bytes and file metadata are different problems: bytes go to object storage, metadata goes to a sharded database. Twelve gigabytes a second must not flow through application servers. And tens of millions of idle devices waiting for changes need a cheap way to wait.

API design

The client splits a file into blocks and hashes each one. It asks which blocks are missing, uploads only those to object storage, then commits.

POST /api/uploads
{ "path": "/docs/plan.pdf", "size": 9437184, "blocks": ["9f2c...", "a71e...", "c03b..."] }
→ 200 OK  { "upload_id": "up_81", "missing": [ { "hash": "c03b...", "url": "https://s3...presigned" } ] }

PUT https://s3...presigned            (raw block bytes, client to S3 directly)

POST /api/uploads/up_81/commit
{ "base_version": 6 }
→ 200 OK        { "file_id": "f_42", "version": 7 }
→ 409 Conflict  { "latest_version": 8 }

GET /api/changes?cursor=ns_9:1042
→ 200 OK  { "changes": [ ... ], "cursor": "ns_9:1057", "has_more": false }

GET /api/changes/wait?cursor=ns_9:1057&timeout=60
→ 200 OK  { "changed": true }

POST /api/shares  { "folder_id": "f_10", "user_id": "u_7", "role": "editor" }

The commit carries base_version, the version the client last saw, so the server can detect a conflict instead of overwriting. The wait call returns only a flag; the client then fetches the changes.

Data model

A file is a list of blocks, and a block is identified by the hash of its content. Metadata lives in a relational store such as RDS PostgreSQL, sharded by namespace - a user's own root or a shared folder. Block bytes live in S3, keyed by hash.

files
  namespace_id    bigint     shard key
  file_id         bigint
  parent_id       bigint     folder
  name            text
  latest_version  int

file_versions
  file_id         bigint
  version         int
  block_hashes    text[]     ordered list, e.g. SHA-256 per block
  size            bigint

blocks
  hash            text       primary key, also the S3 object key
  size            int
  ref_count       int

journal
  namespace_id    bigint     partition key
  seq             bigint     increasing per namespace
  file_id         bigint
  version         int
  op              text       add, edit, move, delete

namespace_members
  namespace_id    bigint
  user_id         bigint
  role            text       owner, editor, viewer

Versions are cheap because they share blocks. Changing one block of a 100 MB file adds one block and one version row; the rest are referenced, not copied.

High-level design

Keep bytes and metadata on separate paths, and keep the expensive work off the request a user waits on.

  • The client watches the local folder, chunks changed files into blocks (4 MB is a typical choice), hashes them, and records what it last synced.
  • The metadata service owns files, versions, blocks, the journal and permissions. It hands out short-lived pre-signed S3 URLs for missing blocks and checks permissions before every one.
  • Block storage is S3. Clients upload and download blocks directly, so the application tier never carries file bytes.
  • On commit, the metadata service writes the new version and a journal entry in one transaction, then publishes a change event to SNS or a pub/sub tier. The notification service wakes up any device waiting on that namespace.

Where it breaks

The simple first design uploads the whole file through the API servers, which write it to storage. At 12 GB/s the API tier becomes the bottleneck, and a small edit to a 2 GB video re-sends 2 GB.

The fix has two parts. Chunk and hash on the client so only new blocks are sent, and send them straight to S3 with pre-signed URLs, so the API servers only handle small metadata calls. The new risk is orphaned blocks: a client uploads blocks, then dies before it commits. Treat a block as live only once a committed version references it, and let a cleanup job delete unreferenced blocks after a grace period long enough to cover in-flight uploads.

Chunking, dedup and delta sync

  • Fixed-size blocks are simple, but inserting one byte near the start shifts every block after it and they all hash differently. Content-defined chunking picks block boundaries from the data itself, so an insert only changes the blocks around it.
  • The block hash is the dedup key. If a hash already exists, the block is not uploaded again. Within one account this makes re-uploads, copies and restores almost free.
  • Dedup across accounts saves more but leaks information: a fast upload reveals that someone else has that exact file. Limit dedup to an account, or make the client prove it holds the bytes.
  • Delta sync falls out of the block model: a download compares the new block list with the local one and fetches only the hashes it doesn't have.

The sync protocol

  • Every change in a namespace gets the next seq number in the journal. A device stores one cursor per namespace: the last seq it has applied.
  • To sync, a device fetches journal entries after its cursor and applies them in order. This works the same after five seconds or five weeks offline.
  • To hear about changes without constant polling, the device holds a long poll (or a WebSocket) to the notification service. It returns when the seq moves past the cursor, or after about 60 seconds.
  • The notification is only a hint. The journal is the source of truth, so a lost notification delays a sync but never loses a change.

Conflicts and offline edits

Two devices edit the same file from version 6. The first commit becomes version 7. The second arrives with base_version 6, which no longer matches, so the server rejects it with a conflict. Don't try to merge arbitrary files. Keep both: the server's version stays as the file, and the client saves its edit as a separate "conflicted copy" next to it, so no one's work is lost. Offline edits use the same path: the client queues changes, syncs the journal on reconnect, then commits each with its base version.

What interviewers look for

  • Separating file bytes (object storage, pre-signed URLs) from metadata (a sharded database), with numbers that show why.
  • Content-hashed blocks, and how they give you dedup, cheap versions and delta sync.
  • A sync design built on an ordered change journal and per-device cursors, with notifications as a hint rather than the source of truth.
  • A clear, safe conflict policy that never silently drops an edit, and a plan for orphaned blocks and permission checks.

Frequently asked questions

Why split files into blocks?

+

Blocks let a client upload only what changed, resume a failed upload, and skip blocks the server already has. They also make versions cheap, because versions share unchanged blocks.

Why upload to S3 with pre-signed URLs instead of through the API?

+

File bytes are most of the traffic. Sending them straight to object storage keeps large transfers off the application servers. The URL is short-lived and scoped to one object, so the server still controls who may write what.

How does a device find out that a file changed?

+

Each namespace has an ordered change journal, and each device keeps a cursor into it. The device holds a long poll or WebSocket to a notification service that wakes it when the journal moves, then it fetches the entries after its cursor.

What happens when two people edit the same file?

+

Each commit includes the version the client started from. The first commit wins; the second is rejected as stale, and the client saves its edit as a conflicted copy beside the file. Nothing is lost.

How does sharing a folder work?

+

A shared folder is its own namespace with a member list and roles. Members see it in their tree, their devices sync its journal, and the metadata service checks the role before issuing any upload or download URL.

Now break one yourself.

The first challenge takes about two minutes. No signup.