“How many servers do we need for ten million users?” sounds like a sizing question, but the user count does not contain enough information to answer it. Ten million registered accounts may produce modest traffic, while a much smaller audience refreshing a live event can overwhelm an unprepared service.

Useful capacity planning turns product behaviour into resource demand, then checks that demand against measured service limits. The answer is a model with assumptions and failure scenarios, not a universal number of requests that one server can supposedly handle.

Introduction

This article develops a practical capacity estimate for a web platform with millions of users. We will convert daily activity into peak traffic, calculate CPU and concurrency requirements, evaluate shared dependencies and account for availability-zone failures. The numbers are illustrative assumptions, not benchmark results from a particular production system.

The aim is to produce a defensible starting fleet size and a clear validation plan. A good estimate explains which inputs are uncertain, which bottleneck currently determines capacity and what evidence would cause the design to change.

We will also distinguish steady-state capacity from burst absorption and recovery capacity. A system that handles an ordinary Tuesday may still fail during a rolling deployment, a cache outage or the return of a large background backlog. Those situations belong in the plan before launch.

Begin with the Service Objective

Define what successful service means before asking how much traffic a machine can process. A server might return ten thousand responses per second while its tail latency is several seconds and many responses are errors. That is not useful capacity for an interactive API promising fast, reliable results.

For this example, assume the main API targets a 250-millisecond p99 response time under its normal request mix, with an agreed error-rate objective. Background reports have a separate completion deadline and should not consume unlimited interactive capacity.

Specify the tested payload sizes, authentication behaviour and downstream calls. A benchmark of a tiny unauthenticated health endpoint cannot determine the capacity of a personalised search or checkout request.

Also identify the failure tolerance. We will require the service to sustain the planned peak after losing one of three availability zones. Later, we will consider whether it must simultaneously tolerate a rolling deployment or whether deployments pause during that incident.

The capacity boundary is the highest sustained useful throughput that meets these conditions with acceptable resource behaviour. It is usually lower than the maximum rate observed immediately before a process crashes.

Convert Registered Users into Daily Work

Suppose the product has ten million registered users, of whom two million are active on a typical day. Each daily active user generates forty API requests across sessions, giving eighty million user-originated requests per day.

Daily requests = daily active users × requests per active user
= 2,000,000 × 40
= 80,000,000

Average requests per second = 80,000,000 / 86,400
≈ 926

The forty-request assumption must represent actual API requests, not screen views if each screen calls several endpoints. Include background refreshes, mobile synchronisation and other client behaviour that reaches the service.

Separate static assets served by a CDN from application requests. A browser may fetch dozens of images without creating the same number of origin API calls. Conversely, one user action can fan out into several internal service calls even when the browser sends only one request.

Daily active users also need segmentation. A casual reader and a bulk-importing business customer may differ by orders of magnitude. If heavy users dominate demand, model their activity separately instead of assuming an average person represents the workload well.

Estimate Peak Demand and State the Time Window

Average traffic is useful for daily cost and data-volume estimates, but it is rarely sufficient for latency-sensitive capacity. Traffic concentrates by time zone, campaign, event and client retry behaviour.

Assume a tenfold peak-to-daily-average ratio for the example, producing roughly 9,260 requests per second. Add a separately justified 10% allowance for health checks, internal control traffic and expected retry overhead, then 15% forecast growth. The result is around 11,700 requests per second, rounded to a planning target of 12,000.

These multipliers are assumptions to replace with telemetry. Do not stack several vaguely named safety factors without explaining what each covers. A peak multiplier derived from data that already includes retries should not receive the same retry allowance again.

State the time window: a one-second spike, a five-minute peak and an hour-long surge require different responses. A queue can absorb a brief asynchronous burst, while a sustained interactive peak needs enough serving capacity or explicit load shedding.

Check correlated peaks. A marketing campaign might bring more users and make each user issue more expensive searches. Multiplying independent average assumptions can miss the fact that volume and per-request cost rise together.

Model the Request Mix

Requests per second are only meaningful when paired with the work each request performs. A cached lookup, a multi-filter search and a generated report can have very different CPU, memory and dependency costs.

Assume the illustrative API mix is 60% simple reads, 20% search, 15% writes and 5% report-related requests. Initial profiling suggests average application CPU costs of one, six, four and twenty milliseconds respectively for those classes.

Weighted CPU time per request
= 0.60 × 1 ms
+ 0.20 × 6 ms
+ 0.15 × 4 ms
+ 0.05 × 20 ms
= 3.4 ms

This weighted average is useful for a first CPU model, but it does not replace per-class latency measurements. A small percentage of expensive requests may dominate tail latency or monopolise a dependency even while average CPU looks comfortable.

Google's SRE chapter on handling overload discusses why raw request counts can be misleading when request costs vary. For our estimate, keeping the request mix explicit lets us see how a new feature or traffic shift changes resource demand.

If report generation is genuinely expensive and delay-tolerant, move it into a bounded worker pool with its own capacity model. The interactive endpoint then accepts a durable job and exposes progress, rather than holding a web request open for minutes.

Calculate a CPU-Based Upper Estimate

An eight-vCPU instance has a nominal budget of eight CPU-seconds per wall-clock second. If we choose a planning utilisation of 60%, its budget for this simplified calculation is 4.8 CPU-seconds per second.

Dividing by 0.0034 CPU-seconds per request gives approximately 1,410 requests per second. This is a modelled CPU estimate, not a guarantee. Runtime overhead, uneven scheduling, garbage collection, operating-system work and other bottlenecks can lower the useful rate.

Do not confuse CPU time with response time. A request can take 100 milliseconds end to end while consuming only a few milliseconds of CPU because it waits for network or storage. Dividing one second by response time does not tell you how many requests a concurrent server can handle.

The utilisation target is also workload-specific. A service with predictable CPU work and robust admission control may operate differently from one with bursty allocation and tight tail-latency requirements. Choose the target from the latency curve and failure behaviour, not from an inherited rule that every server should stay below one particular percentage.

Different processor families and virtualisation conditions can change the measured cost. “Eight vCPUs” is not a universal unit of application throughput. Revalidate when changing instance type, runtime version or a substantial part of the request implementation.

Measure Safe Per-Instance Capacity

Now test one production-like instance with the representative mix and dependencies. Increase offered load in controlled steps, hold each step long enough to observe steady behaviour and record latency, errors and resource consumption.

Suppose this illustrative test finds that 1,200 requests per second meets the service objective with stable memory and acceptable dependency pressure. At 1,500, p99 latency breaches the target even though the process continues responding. Use 1,200 as the safe planning rate for this environment.

Include authentication, serialisation, logging and realistic payloads. Disable neither validation nor security features merely to produce a favourable headline throughput. Warm-up behaviour and connection establishment should be measured separately because new instances experience them during scaling and recovery.

Use several repetitions and realistic data distributions. A uniformly generated dataset can hide one large tenant, an expensive query range or a cache hotspot. The benchmark should include the cases the production service will actually encounter.

For .NET applications, tools such as dotnet-counters can expose runtime measurements during testing. Combine runtime evidence with operating-system, database and tracing data so that the limiting resource is identified rather than inferred from one CPU chart.

Turn Per-Instance Capacity into a Fleet Size

At a planning peak of 12,000 requests per second and a safe rate of 1,200 per instance, the service needs ten healthy serving instances. That is the workload requirement before redundancy and placement constraints.

Healthy instances required = ceiling(12,000 / 1,200)
= 10

With three equally sized availability zones and a requirement to survive losing one, deploy five instances per zone. Fifteen total instances leave ten after one zone is lost, matching the modelled peak requirement.

This assumes traffic can move to the surviving zones, shared dependencies also survive and the per-instance measurement remains valid under that topology. Cross-zone network traffic, cache locality and database failover can change performance during the incident.

If the service must also lose one instance for deployment while a zone is unavailable, fifteen is insufficient. One example placement of seventeen instances is six, six and five; losing a six-instance zone leaves eleven, and draining one more leaves ten. Whether that combined scenario is required is a deliberate availability and cost decision.

Round and place capacity carefully. A formula using a surviving fraction is a useful estimate, but actual integer placement can leave the largest zone holding more capacity than expected. Validate the worst allowed placement, not only the average distribution.

Check Concurrency and Memory

Throughput and latency imply concurrent work. At 12,000 requests per second and an average end-to-end time of 100 milliseconds, the system has about 1,200 requests in flight under stable conditions. This is an average relationship, not a bound on bursts or tail behaviour.

If latency rises to one second while the arrival rate stays unchanged, in-flight work can grow towards 12,000. Each request may retain buffers, authentication state and objects while waiting. An I/O slowdown can therefore create memory pressure even when CPU is not initially saturated.

Estimate retained memory per active request, then add caches, runtime baseline, connection buffers and background work. Large request bodies should be streamed or limited appropriately; one average allocation number can hide a small class of requests that consumes most memory.

Bound concurrent work and queues. Unlimited acceptance does not improve capacity when the service is already waiting on a saturated dependency. It can instead increase garbage collection, timeouts and retries until useful throughput falls.

In managed runtimes, distinguish allocation rate from retained heap size. A small live heap with a very high allocation rate can still consume substantial CPU through garbage collection. Both measurements help explain why a CPU model changes as throughput increases.

Size the Database from Actual Operations

One API request does not equal one database operation. A write endpoint may read permissions, update several rows, insert an audit record and commit a transaction. A search may call a separate search engine, while a cached read may never reach the database.

At 12,000 API requests per second, a 15% write share produces 1,800 write requests per second. If each performs three relevant database operations, that path alone creates roughly 5,400 operations per second. The operations differ in cost, so do not turn that count directly into a database size without measurement.

Model transaction duration, rows touched, index maintenance, log generation and lock contention. A single hot inventory row can limit throughput even when database CPU and aggregate IOPS appear low. Adding read replicas will not accelerate that serial write boundary.

Connection budgets need fleet-wide arithmetic. Fifteen application instances with a maximum pool of one hundred connections can potentially demand 1,500 connections, before workers and operational tools are included. A large pool can move queuing into the database rather than increasing useful work.

Test the database with the complete expected workload and failure topology. Application fleet scaling is only safe if shared state can handle the increased demand. A stateless tier sized correctly above an undersized database remains an undersized system.

Model Cache Hits and Cache Failure

Suppose simple reads account for 7,200 requests per second at peak and 90% are served by a cache. The database sees about 720 misses from that path. If the hit rate drops to 80%, misses double to 1,440 even though user traffic has not changed.

That sensitivity matters during deployment, eviction, a cache restart or a new feature that uses different keys. Planning solely for the normal hit rate can make a routine cache event look like an unexpected traffic attack on the database.

Define the supported degraded behaviour. Can stale values be served for selected data? Can misses be coalesced so many concurrent requests do not regenerate the same item? Which requests should be rejected if the origin cannot absorb the miss load?

Warm caches gradually and avoid synchronised expiry of a large key population. New instances may also have empty local caches even when the shared cache is healthy, changing their initial CPU and dependency costs.

Do not assume the entire database must always sustain every cacheable request at full peak. That may be prohibitively expensive. Instead, state the failure policy explicitly and test the combination of reserve capacity, coalescing, bounded staleness and admission control that preserves the most important product behaviour.

Include Network and Payload Size

At 12,000 responses per second and an average response payload of twenty kilobytes, application egress is roughly 240 megabytes per second, or 1.92 gigabits per second before protocol overhead. Compression, CDN use and response-size distributions change the practical result.

Check both aggregate and per-instance bandwidth. Load distribution may be uneven, and some instance types have burstable network limits or different packet-processing characteristics. Small packets can stress connection and packet rates even when total bytes look modest.

Internal fan-out adds traffic that is absent from browser analytics. One API request might call five services, each returning data that is later reduced to a small client response. Distributed tracing can reveal this amplification and its cross-zone cost.

TLS handshakes and connection churn also consume resources. Reusing connections can help, but long-lived connections affect load distribution and memory. A fleet with sticky sessions or persistent streams may not rebalance instantly when new instances arrive.

Large media should normally use storage and CDN paths suited to that workload. Routing every download through the general API tier can make the web-server estimate depend more on file transfer than on application processing.

Plan Storage Growth and Retention

Storage capacity begins with durable business records, but it also includes indexes, replicas, logs, backups and temporary operational space. Avoid presenting only raw row bytes as the final requirement.

For a separate illustrative storage assumption, suppose the platform creates five million retained records per day at one kilobyte each. That is about five gigabytes of raw daily growth. Indexes and record overhead might increase the live footprint substantially, and replicated copies multiply physical storage further.

Use measured sizes once a representative dataset exists. Text lengths, JSON fields, compression and index choices can make an assumed one-kilobyte row inaccurate. Updates may also generate log and version-retention work without increasing the visible row count.

Define retention by data class. Operational traces, temporary exports and permanent customer records rarely need identical lifetimes. Deletion or archival pipelines themselves need capacity and monitoring; a retention policy that is never executed does not control growth.

Reserve headroom for index builds, compaction, backfills and replica catch-up. A database that fits today's live data with only a few percent free can still fail during an ordinary maintenance operation. Capacity planning includes the work required to keep the system healthy.

Account for Background Work and Recovery

Backups, analytics, search reindexing and reconciliation often run outside the interactive request path but still compete for CPU, I/O and database connections. Schedule or isolate them according to their deadlines and the headroom available.

Queues create another obligation: recovery capacity. If workers fall behind during an outage, simply restoring a completion rate equal to the new arrival rate will never drain the backlog. The service needs spare throughput, reduced arrivals or an explicit expiry policy.

For example, a backlog of 360,000 jobs with 800 new arrivals per second and capacity for 1,000 completions per second takes about thirty minutes to drain at the net rate of 200 per second. Dividing the backlog by total capacity would underestimate recovery time.

Replays should be rate-limited and observable. Restarting every failed job at once can overload the dependency that just recovered. Reserve capacity for current critical work while steadily paying down the historical backlog.

Include restoration exercises. Rebuilding a cache, restoring a database backup or recreating a search index can demand very different resources from steady-state serving. Recovery objectives are credible only when the infrastructure can perform that extra work within the required time.

Use Autoscaling as a Control Loop

Autoscaling responds to a signal after demand or resource use changes. It takes time to observe the signal, decide, obtain capacity, start the application and warm dependencies. A sudden burst can arrive faster than the loop can react.

Keep enough baseline capacity for predictable demand and use scheduled scaling where events are known in advance. Reactive scaling handles additional variation, with limits that protect shared dependencies.

Kubernetes Horizontal Pod Autoscaling documentation describes resource and custom-metric scaling behaviour. The relevant lesson is to choose a signal tied to the workload and understand the controller's timing; adding replicas is not an instantaneous response to a graph crossing a line.

CPU is useful for CPU-bound work. Queue age or active concurrency may be more informative for a worker or I/O-bound service, provided the downstream system can accept more parallelism. Scaling on queue depth alone can add idle or throttled workers when the true bottleneck is elsewhere.

Test scale-in as well as scale-out. Draining connections, completing active work and respecting minimum failure capacity prevent cost optimisation from creating request loss or a redelivery storm. Avoid oscillation by using appropriate stabilisation and separate decision thresholds.

Validate with the Right Load Model

A fixed number of virtual users often produces a closed workload: each user waits for a response before sending more work. When the system slows, the generator may send fewer requests, hiding how an external arrival stream would continue building pressure.

An arrival-rate model schedules new work independently of previous completion, subject to the generator having enough resources. It is useful for testing a specified offered rate and observing overload honestly. Grafana's k6 explanation of open and closed models describes this distinction.

Monitor the load generator itself. If it cannot create connections or execute scenarios quickly enough, an apparent server ceiling may actually be a client ceiling. Use multiple generators where justified and record achieved offered load separately from completed responses.

Test normal peak, zone failure, cold caches, deployment overlap and dependency slowdown. The same nominal request rate can produce very different results in those scenarios. Keep the request mix and dataset version with every benchmark so comparisons remain meaningful.

Check recovery after the test's peak ends. Memory should stabilise, queues should drain and tail latency should return to normal. A service that keeps accepting requests while accumulating hours of work has not demonstrated sustainable capacity.

Run a Sensitivity Check before Committing Spend

The most useful next calculation is often a change to an assumption. Suppose report-related requests rise from 5% to 10% of traffic while simple reads fall from 60% to 55%. Search and write shares remain unchanged. Using the earlier illustrative CPU costs, weighted CPU demand rises from 3.4 to 4.35 milliseconds per request.

That is roughly a 28% increase in application CPU work with no increase in request count. If CPU remained the limiting resource and other conditions were unchanged, scaling the 1,200-request benchmark proportionally would suggest about 940 requests per second per instance. This is a projection to test, not a replacement benchmark.

At that projected rate, the planned peak would require thirteen healthy instances instead of ten. A fleet sized precisely for the original mix could therefore miss its objective during a feature launch even though the traffic dashboard never exceeded the expected requests per second.

Compare this sensitivity with an optimisation. If profiling identifies avoidable report serialisation work, reducing that cost may recover more useful capacity than adding web instances. If the work belongs in a separate asynchronous service, moving it can also make interactive capacity more predictable.

Perform similar checks for cache misses, response size and tenant concentration. Use low, expected and high scenarios with reasons for their bounds. Do not manufacture precision by reporting a fleet requirement to several decimal places when the peak multiplier itself is uncertain by a factor of two.

This analysis guides measurement priorities. Spend effort validating assumptions that materially change the design, rather than refining a minor storage estimate while the request mix remains unknown. A capacity model earns its value by identifying the next useful experiment and the point at which a different architecture or resource allocation becomes necessary.

Present the Estimate as a Range with Decision Points

For our illustrative API, ten healthy instances meet the planned 12,000-request-per-second workload at a measured safe rate of 1,200 each. Fifteen across three zones meet the single-zone-loss assumption. Seventeen in an appropriate placement cover the additional example of draining one instance after that loss.

Those are application-serving counts, not the total number of machines in the system. Databases, caches, brokers, search services and background workers have their own measurements and redundancy requirements. Managed services still consume capacity even when the team does not count their hosts directly.

Identify the most sensitive assumptions: daily activity, peak multiplier, request mix, cache hit rate and per-instance safe throughput. If a small change in one input doubles the database load, prioritise measuring and protecting that input.

Keep a capacity record with observed peaks, tested limits, expected growth and the lead time for adding resources. Google's SRE workbook chapter on managing load provides broader operational context for connecting provisioning and load-management decisions.

Revisit the model after major releases and workload shifts. Capacity planning is not a one-off interview calculation that becomes permanently true. Its value is that new evidence can be inserted into a clear model and translated into an action before customers experience the limit.

Summary

User count is the beginning of capacity planning, not the answer. Convert activity into a time-based request profile, describe the request mix and measure the useful throughput that meets the service objective on representative infrastructure.

Size the healthy fleet, then apply explicit failure and placement requirements. Check concurrency, memory, database operations, cache misses, network traffic and storage independently, because the narrowest shared bottleneck determines the system's real capacity.

The best estimate includes its own validation plan. It explains the assumptions, tests degraded scenarios and reserves room for recovery as well as ordinary traffic. That turns “how many servers?” into a decision the team can defend, measure and update as the product grows.