System Design Sim

Learn

How many users can one server handle? A worked estimate

"How many users can one server handle?" has no single answer, but it has a method. Turn users into requests per second, compare that with what one server can do, and leave room for traffic peaks and a server failing. Here is the method with real AWS sizes and prices.

· 3 min read

How many servers do you need?

1M users make about 231 req/s on average and 463 req/s at peak. The cheapest fit is 3 × m7g.medium for about $89/month, sized to 70% CPU with one spare.

Servers needed for peak traffic by instance size
SizevCPU · RAMHandles eachServersPer month
m7g.medium1 · 4 GiB~500 req/s3$89
m7g.large2 · 8 GiB~1,000 req/s2$119
m7g.xlarge4 · 16 GiB~2,000 req/s2$238
m7g.2xlarge8 · 32 GiB~4,000 req/s2$477
c7g.large2 · 4 GiB~1,000 req/s2$106
c7g.xlarge4 · 8 GiB~2,000 req/s2$212
c7g.2xlarge8 · 16 GiB~4,000 req/s2$423

Approximate AWS on-demand prices for us-east-1, read once in October 2026, before free tiers. They may be out of date; check the AWS pricing pages before you budget.

Step 1: turn users into requests per second

Servers don't see users, they see requests. Convert with three numbers:

  • Daily active users (DAU): people who use the product on a given day.
  • Requests per user per day: every page load, API call and background refresh counts. 10-50 is common for an app.
  • Peak factor: the busiest hour is usually 2-3× the daily average.

Peak requests/s = DAU × requests per user ÷ 86,400 × peak factor. One million users at 20 requests a day is 20 million requests, about 231 per second on average and 463 at a 2× peak. A million users sounds like a lot; 463 requests per second is not.

Step 2: what one server can do

It depends entirely on the work per request. Published measurements range widely: a 1-vCPU server running Node.js measured about 1,300 requests/s on a simple JSON endpoint and about 1,050 with a SQLite query (measured here), while heavily tuned servers returning static JSON reach hundreds of thousands.

A useful planning number for a typical API endpoint that makes one cache or database call is about 500 requests/s per vCPU, so a 2-vCPU m7g.large handles about 1,000. Treat it as a starting point and replace it with a load test of your own service as soon as you can.

Step 3: leave headroom

Never plan to run servers at 100%. Two reasons:

  • Latency. As a server gets busy, requests wait for each other. Past roughly 70-80% utilization the queue grows fast and p99 latency climbs steeply, long before the server is "full".
  • Failures. When one server dies, the others take its traffic. Run one more than you need (N+1), behind a load balancer that routes around the dead one.

The calculator above and the table below size for 70% CPU at peak, plus one spare.

Worked examples

At 20 requests per user per day and a 2× peak:

Daily usersAveragePeakCheapest fleetPer month
10K2 req/s5 req/s2 × m7g.medium$60
100K23 req/s46 req/s2 × m7g.medium$60
1M231 req/s463 req/s3 × m7g.medium$89
10M2,315 req/s4,630 req/s8 × c7g.large$423
100M23,148 req/s46,296 req/s68 × c7g.large$3,599

The API tier is rarely what limits you, and it's the easiest part to scale: add servers. The database behind it is harder.

What usually breaks first

Stateless API servers scale out almost linearly. The shared things behind them don't: a single database primary, a cache that's too small, or a downstream service with its own limits. Once you know your peak requests per second, size those next:

Frequently asked questions

How many requests per second can one server handle?

It depends on the work each request does. A useful rule of thumb for a typical JSON API that makes one cache or database call is about 500 requests per second per vCPU, so a 2-vCPU m7g.large handles about 1,000. A trivial endpoint goes much higher; one that renders pages or calls slow services goes much lower. Measure your own with a load test.

How many servers do I need for 1 million daily users?

At 20 requests per user per day, 1 million users make 20 million requests a day: about 230 per second on average and about 460 at a 2× peak. One m7g.large handles that, but run at least two behind a load balancer so a single failure doesn't take the service down.

Why not run servers close to 100% CPU?

Because waiting time grows sharply as a server gets busy: past roughly 70-80% utilization, requests start queueing and latency climbs steeply. You also need spare capacity so that when one server dies, the others can absorb its traffic.

One system design problem a week

A real scenario, its constraints, and where the obvious design breaks.

One email a week. Unsubscribe any time. Privacy

Now break one yourself.

The first challenge takes about two minutes. No signup.