User count is not a workload model

Ten thousand registered users do not determine server count. Model operations: arrivals per second, work per request and bursts. One customer exporting a large dataset can cost more than hundreds of people opening a page.

We split quick reads, writes, expensive computations and background jobs. Each class gets its own latency budget and concurrency limit. One aggregate average lets cheap reads hide slow reports, and the load test can look much healthier than the workload we actually need to run.

In an illustrative stable workload, 200 requests per second and 0.2 seconds of average time in the system imply roughly 40 concurrent requests. If average time grows to two seconds at the same rate, concurrency rises to about 400. This relates averages in a stable system; it does not calculate p99 or a required connection-pool size.

We model a concrete hour of product use: order lists, occasional updates and large reports. Frequency and resource cost are separate. Reports may be one percent of requests but consume most CPU or disk reads. Treating request share as cost share can make the cheapest operations dominate the plan.

Bursts have different shapes. A synchronized login after an email campaign stresses cold caches and rapid concurrency growth. A day-long rise tests sustained throughput. We reproduce the shape relevant to the product rather than run one maximum-RPS number and assume it describes both conditions.

The response budget applies across the path. A one-second promise cannot give each of five sequential calls a one-second timeout. Network time, serialization and queues consume part of it too. The budget exposes which dependency already uses the promise, even when each service considers itself quick.

Work size needs bounds. A ten-row list and a year-long export might use the same endpoint but require very different capacity. Limiting concurrency without limiting input size gives weak control. Heavy operations need result limits, execution budgets and explicit rules for background processing.

Data growth belongs in the model. An empty database does not represent next year's workload. Large accounts, indexes and popular-object distribution change request cost. We record volumes and skew alongside the test results, so later comparisons explain why the same RPS no longer fits.

Find the service bottleneck: processing or waiting?

A request can spend 30 ms computing and 900 ms waiting for a database connection. Application CPU stays available. Extra instances create more contenders without speeding up the database. Measure pool waits, queues, locks and external calls alongside total latency.

Increase load gradually and find where throughput stops growing while unfinished work accumulates. An unlimited queue postpones rejection but adds waiting, memory consumption and work for clients who have already left.

We set the operating limit below the point where response time starts falling apart. Headroom covers ordinary bursts and losing part of our capacity. A rule like "keep CPU below 70%" cannot tell us that boundary: it depends on the workload and where requests wait.

We compare incoming work, successful completions and work in progress. This reveals when the service stops keeping up. Errors alone are late evidence: users may still be waiting while resources are already occupied by operations heading toward timeouts. Detecting accumulation gives us a chance to act earlier.

One slow trace identifies a path, but does not establish its prevalence. We check how many requests wait there and compare with aggregate signals. Otherwise a day of investigation can optimize an unusual tail while missing a much more common pool wait that controls overall behavior.

Pool wait and SQL execution remain separate measurements. A query running for fifteen milliseconds after a five-hundred-millisecond connection wait does not necessarily need a faster SQL plan. Excess concurrency, leaked connections or transactions containing external calls may be responsible. Measuring only database execution would miss them.

We change one variable during an experiment. For example, limit simultaneous reports while keeping resources constant and rerun the same traffic. Faster interactive requests then provide evidence of contention. Changing server, pool and several parameters together makes an attractive result harder to explain or repeat.

After a fix, we measure the whole chain again. The bottleneck usually moves: CPU gives way to database or upstream limits. That is expected. The mistake is turning the first successful scale-out into a permanent rule to add instances without checking where the new capacity stops helping.

Each process adds its own dependency demand

A pool allowing 20 connections per process can produce 80 connections across four processes and 400 across twenty. If the database supports fewer, application autoscaling hits a shared limit. Budget connections across applications, background jobs and administration.

Check local state too. A file on one instance does not appear on another; an in-memory session can disappear during routing; a scheduled task can run once per process. Each requires an explicit ownership or storage decision, not merely a load balancer.

We do not automatically enlarge the pool with request volume. More connections can mean more competing database work and worse response time. We find a range that completes useful work without excessive pressure. Different operation classes may need separate budgets to protect the interactive path.

Files need shared storage or an explicitly local use case. Sessions need verification by any instance. Scheduled jobs need ownership or safe duplicates. Pinning work to one instance can be a temporary choice, but its failure becomes a scenario to handle rather than an assumption the load balancer resolves.

A load balancer distributes requests, not necessarily cost. One client issues expensive operations, others cheap reads; connections live for different durations and processes warm at different speeds. We inspect per-instance latency and work distribution so a fleet average does not hide a single overloaded process.

Scale-down is part of scaling. We stop admitting new work, allow current operations to finish within a limit and handle unfinished tasks explicitly. Queue return or result reconciliation may be needed. Otherwise ordinary deployments and capacity reduction will repeatedly lose work despite the fleet appearing healthy.

The full instance cost includes connections, metrics, local cache and warmup reads. Starting several processes together can briefly overload a shared dependency. Gradual traffic admission and a fleet-wide budget keep capacity expansion from becoming an incident of its own precisely when the system has little spare headroom.

Bound the work before resources are exhausted

Limit concurrency, queue length and waiting time. Rejecting some work promptly can preserve useful throughput instead of letting all requests expire. Priorities can protect checkout while a heavy export temporarily waits or fails.

Retries without a shared budget can amplify an outage. Three call layers making up to three attempts each can turn one request into 27 calls to the final dependency. We choose which layer owns the retry, cap attempts and use backoff with jitter so clients do not retry together. Timing out only tells us that we did not receive an answer.

For repeated writes, use an operation identifier and durable duplicate handling. An in-memory check disappears when the process fails. Recording the result and returning it again must remain consistent with the actual data change.

We reject work where rejection is still cheap. Once a queue consumes memory and requests open transactions, a late refusal cannot recover every cost. Input limits, concurrency control and admission checks belong before the expensive stage. Their location matters as much as the limit itself.

Backpressure requires upstream behavior to change when the next stage lacks capacity. A durable queue only postpones the problem if producers keep outrunning workers. We need a maximum useful job age and a decision for stale tasks, not just a larger place to store the backlog.

Retries help when another attempt can change the result. Repeating invalid input is useless. A transient network failure may justify a repeat if the operation is safe and time remains. The budget follows the total user deadline, rather than resetting an independent timeout on every attempt.

Jitter prevents clients from returning together. Telling every client to wait one second simply creates another synchronized arrival one second later. Random delay still does not create capacity. Sustained overload needs reduced input, cheaper work or a resource increase at the actual bottleneck.

Cancellation deserves a path too. A customer leaves while the server continues building a report and calling dependencies. Where safe, we propagate cancellation and free resources. Once an operation changes data, cancellation of the client's wait must not leave the business result undefined. Completion and later retrieval need separate rules.

Autoscaling cannot arrive instantly

An instance needs startup, warmup and traffic admission. A shorter burst can overwhelm the service before reactive capacity arrives. Predictable events may need capacity in advance; unpredictable ones need limits and defined overload behavior.

CPU is not an adequate signal for every workload. Queue workers need arrival rate, completion rate and oldest-job age. A thousand quick jobs and a thousand minute-long jobs require different responses. Stop adding workers when the downstream dependency becomes the limit.

We measure the entire reaction delay: metric collection, controller decision, process startup, warmup and readiness for useful traffic. A quick container start does not establish quick usable capacity. A service with expensive warmup may need a permanent reserve despite having a reactive controller.

For a queue, we estimate completion time as well as backlog. When input stops, job count divided by completion rate gives a rough drain time. With continuing input, the difference between arrival and completion rates matters. Unequal job classes need separate estimates or quick jobs conceal a slow batch.

Adding and removing capacity deserve different caution. Rapid addition may help; rapid removal after a short dip can recreate the queue. Thresholds, delays and minimum instance count account for traffic shape, warmup costs and recovery expectations. They should be validated against the workload rather than chosen from a generic example.

We check the upper bound before relying on it. Quotas, addresses, connection budgets and database capacity can stop expansion earlier than the desired process count. Reaching that limit needs a visible signal. Enabling autoscaling does not establish that more capacity is actually available during a peak.

The control signal can disappear, arrive late or change meaning after a release. We define controller behavior under missing or bad data and bound erroneous scaling. Automation reduces manual work when these failures are understood; otherwise it adds a decision-maker whose mistakes must be diagnosed during the incident.

Test capacity together with failure

Use representative data sizes, operation distribution and client retry behavior. Remove an instance or slow a dependency during the test. Capacity that exists only when every resource is available does not provide headroom for that failure.

After the change, we compare successful throughput, tail latency, cost and recovery. We record the test conditions and the limit we found. The next step might be a query fix or a pool setting rather than another instance. Useful work completed by the service is the outcome we care about.

The generator must preserve the intended arrival pattern as responses slow. If each virtual client waits for completion before sending again, longer latency lowers offered traffic. A beautiful report may then exist because the test stopped producing the real peak. We choose the generation model deliberately.

Generator and network limits are checked separately. An exhausted load machine sets its own ceiling. One account and one set of keys can create unusual contention or an unrealistically warm cache. Test conditions need to be repeatable and explainable, including the dataset and how traffic is distributed.

Failure scenarios are specific: one missing instance, an upstream five times slower or intermittent connection loss. We observe preserved operations and recovery duration. "Handles failures" has no testable meaning without the conditions and guarantees, because different failures affect different parts of the path.

We measure cost per successful useful operation. Doubling resources for five percent more throughput may be acceptable briefly but expensive as a permanent solution. Dependencies and operational effort belong in the cost. This keeps scaling connected to the business result rather than the server count.

The result becomes operating guidance: a safe workload range, reserve, controller conditions and actions outside the range. We recheck after major data or request-path changes. Capacity has a date and test conditions; it is not an eternal property of the service that survives every release unchanged.

Working through an order peak

Before a promotion, normal traffic fits but peak tail latency rises. Traces show short SQL execution and long pool waits; concurrent reports occupy the database. Extra instances without a new connection budget increase contention. The first experiment limits reports and compares the interactive path under unchanged traffic.

If checkout improves, we separate foreground and background budgets. Admission has controlled rejection, queues have maximum useful age and retries share a deadline. A lost response must not create a second order. We retain headroom for an instance failure and prepare warm capacity before the predictable peak.

The next test preserves arrivals despite slower responses and uses large accounts and a cold start. We slow an upstream and remove a process. Useful completions, tail latency and recovery matter together. Faster rejection with fewer completed orders does not establish a better product result.

We carry the measured boundary into operations with workload, dataset, dependency conditions and reserve. Autoscaling stops at a visible shared limit. Later growth is evaluated against the new bottleneck. Capacity planning becomes a sequence of tested changes rather than a promise that containers can expand without affecting the rest of the system.

Back to articles