Calcylator
Performance

Requests per second:
sizing capacity for an API

Convert logs and traffic estimates into load figures, in-flight requests and bandwidth you can plan servers around.

Calcylator Editorial Team

Updated · 5 min read

Counting requests per second from logs or totals

Requests per second (RPS, sometimes QPS) is the arrival rate of requests over a period. You can read it from a monitoring graph or compute it from a total: count the requests in a window and divide by the window length in seconds.

Request rate =total requests ÷ seconds
total requests:
count in the window
seconds:
length of the window
  • Requests in the hour

    1,800,000

  • Window

    3,600 s

Average load

500 requests per second

1,800,000 ÷ 3,600 = 500.

Averages hide peaks. Traffic over a day is rarely flat, so plan from the busiest minute or second you expect, not the daily mean. A peak-to-average ratio of 3 is common for consumer products, but your own data should settle it.

Latency turns rate into concurrency

How many requests are being handled at the same instant depends on how long each one takes. This is Little's Law, and it holds for any stable queue: the average number in the system equals arrival rate times average time in the system.

Requests in flight =RPS × average latency (seconds)
RPS:
arrival rate
latency:
mean time to serve one request, in seconds
  • Rate

    500 RPS

  • Average latency

    40 ms = 0.04 s

In flight

20 concurrent requests

500 × 0.04 = 20. If latency creeps to 200 ms, the same traffic holds 100 requests open at once.

That second number is why a latency regression can exhaust a thread pool or database connection limit even though traffic did not change.

What one worker can handle

A worker that processes one request at a time at 50 ms each can complete 1 ÷ 0.05 = 20 requests per second. Eight such workers give 160 RPS at best. Real capacity is lower, because you should not run at 100 percent utilization; queues grow sharply as a service nears saturation.

  • Plan to run at 60 to 70 percent of the measured maximum so spikes do not cause timeouts.
  • Measure capacity with a load test of the real endpoint, because database calls and payload size dominate.
  • Async or multi-threaded servers handle many requests per worker, so the same sum uses concurrency slots, not processes.

Planning for the peak, with headroom

Suppose the daily average is 500 RPS, the busiest period runs at three times that, and each worker manages 20 RPS. Plan from the average and you would choose 25 workers, which collapses at the first busy hour.

  • Average rate

    500 RPS

  • Peak factor

    3, so 1,500 RPS

  • One worker

    20 RPS

  • Target utilisation

    65%

Workers needed

116

1,500 ÷ (20 × 0.65) = 115.4, rounded up to 116.

The gap between 25 and 116 is why the peak figure and a utilisation target belong in every capacity estimate. Autoscaling softens it, but it still needs a sensible per-worker figure to work from.

From RPS to bandwidth

API throughput is often discussed as data volume. Multiply request rate by the average response size, then by 8 for bits.

Data rate =RPS × average response size
response size:
bytes per response, headers included
Multiply bytes per second by 8 to get bits per second.
  • Rate

    500 RPS

  • Response size

    12 KB

Outbound data

6,000 KB/s = 6 MB/s = 48 Mbps

500 × 12 = 6,000 KB per second; × 8 ÷ 1,000 = 48 Mbps.

Check this against your network link and any per-GB egress cost. Compression and caching reduce it, and a tool for API throughput does the multiplication consistently across units.

Rate limits, quotas and bursts

Many third-party APIs publish limits such as a number of requests per minute. Convert them to the same unit before comparing: 600 per minute is 10 per second. If your job needs 50,000 calls, that is 5,000 seconds, or about 83 minutes, at the cap.

When you cannot meet the rate, the options are batching, caching, a higher plan, or a more patient schedule. All of them start from the arithmetic above.

Reading a load test properly

A load test gives you the number to put into the formulas, but only if you read it carefully. Ramp the load in steps and record throughput, error rate and a high percentile of latency, such as p95, at each step. Throughput rises with load until a resource saturates, then flattens while latency climbs and errors begin.

Closed-loop tools that run a fixed number of clients obey the same law: with 64 clients and 80 ms average latency, the most you can see is 64 ÷ 0.08 = 800 RPS. If latency doubles, throughput halves, so a low reading may reflect a slow dependency, not a slow server.

  • Test from a machine that is not itself the bottleneck.
  • Use realistic data and request mixes, because cached and empty paths flatter the result.
  • Quote the capacity at which latency and error targets are still met, not the highest number achieved.

Estimating load before you have traffic

For a product that is not live yet, work from users. Take 200,000 daily active users who make 25 requests each per day. That is 5,000,000 requests a day, and spread across 86,400 seconds it is about 58 RPS on average.

  • Daily active users

    200,000

  • Requests per user per day

    25

  • Peak factor

    5 (a guess to validate)

Peak load

≈ 289 RPS

200,000 × 25 = 5,000,000 per day; ÷ 86,400 = 57.9 RPS average; × 5 = 289 RPS.

Treat each term as an assumption to be replaced with measurements. A single screen may fire several calls, polling and retries add traffic, and a mobile app launched at the same hour by a push notification produces a spike far above the daily pattern. Note too that RPS, bandwidth and concurrency are different things. A hundred slow downloads can saturate a link at a trivial request rate, while a thousand tiny calls might barely register on the network yet exhaust a connection pool.

Write the estimate down with its assumptions and revisit it after the first real week of data.

Finally, keep the units honest in reports. Requests per second, requests per minute and requests per day are all common in vendor documents and dashboards, and a slip of sixty or eighty-six thousand between them is easy to make. Put the unit in the column header, convert everything to one before comparing, and keep the window and percentile next to any latency value so that two people reading the same chart reach the same conclusion about capacity.

Common questions

How do you calculate requests per second?

Divide the number of requests by the time window in seconds. For instance, 1,800,000 requests over one hour is 1,800,000 ÷ 3,600 = 500 requests per second. Use the busiest window you care about, not just a daily average.

How many requests per second can a server handle?

It depends on the work done per request. A worker that takes 50 ms per request serves about 20 per second, and cores or workers multiply that. Load-test the real endpoint to get a reliable number.

What is the relationship between latency and concurrency?

Concurrent requests equal arrival rate multiplied by average latency in seconds. At 500 RPS and 40 ms, about 20 requests are in flight. Latency of 200 ms at the same rate holds 100 requests open.

How do I convert RPS to bandwidth?

Multiply RPS by average response size, then convert bytes to bits. 500 RPS with 12 KB responses is 6 MB per second, or 48 Mbps. Include headers and encoding overhead for a more accurate figure.

What is a good peak-to-average ratio for traffic?

It varies by product, and your own metrics beat a rule. Consumer apps often see a peak two to five times the daily average. Size for the peak and keep headroom of 30 percent or more.

Was this guide helpful?

Continue reading

View all blogs