API error rate:
from raw counts to an error budget
How to count errors honestly, convert an SLO into a number of allowed failures, and tell when a bad day is actually a problem.
Calcylator Editorial Team
Updated · 4 min read
Deciding what counts as a failure
The arithmetic of an error rate takes one line: failures divided by requests. The hard part is agreeing on both numbers. A service that returns HTTP 200 with a body full of errors can look perfect on a dashboard built from status codes alone.
GraphQL shows this clearly. A single endpoint nearly always answers 200, and problems arrive in an errors array inside the response. Counting HTTP failures only would report a rate near zero even when many queries come back broken, so count operations that returned a non-empty errors field, and decide separately whether partial data counts as a failure.
Server faults and client mistakes also behave differently. A burst of 404 or 401 responses usually reflects callers, not your system. Many teams report 5xx responses (plus timeouts) as the service error rate, and track 4xx separately as a signal of integration problems.
Decide, too, how to treat requests that never completed. A call that timed out on the client may not appear in server logs at all, and a request cancelled by the user is not your fault. Write the rule down once so that every dashboard and every report counts the same way.
The formula and a daily example
- failed requests:
- 5xx responses, timeouts and error payloads you decided to count
- total requests:
- all requests in the same window and from the same source
Requests in the day
3,000,000
Failed requests
4,200
Division
4,200 ÷ 3,000,000
Daily error rate
0.14%
4,200 ÷ 3,000,000 = 0.0014, which is 0.14% of requests failing, or about 1 in 714.
Use the same window and the same source for both counts. Mixing failures from the load balancer with requests from the application log is a common way to publish a wrong rate.
Turning an SLO into an error budget
A service level objective states the success rate you promise over a period. Everything between that promise and 100% is the budget: the failures you can absorb before the objective is broken.
- SLO:
- target success rate, such as 99.9
- total requests:
- expected traffic over the whole SLO window
SLO
99.9% success over 30 days
Expected traffic
50,000,000 requests
Budget fraction
1 − 0.999 = 0.001
Failures allowed in 30 days
50,000 requests
0.001 × 50,000,000 = 50,000 failed requests before the objective is missed.
The budget is a spending limit. A risky release, a migration or a chaotic week of incidents draws it down, and when it runs low the team slows feature work in favour of reliability.
Burn rate: is today's bad day a real problem?
Raw error rate does not say whether you are in trouble, because the same 0.14% is harmless against a 99% target and costly against 99.99%. Burn rate compares what you are seeing with what the objective allows.
For the 3,000,000-request day, the observed rate was 0.14% against an allowed 0.1%, so the burn rate is 1.4. If every day looked like this one, the 30-day budget would be gone after about 21.4 days, roughly nine days before the window ends.
Why you should break the rate down by operation
A single global error rate behaves like an average, and averages flatter. Suppose an API serves three kinds of call in a day. The numbers below are made up, but the pattern is common: the busiest call is healthy, and the rare, valuable one is not.
| Operation | Requests | Failures | Error rate |
|---|---|---|---|
| Product search | 2,400,000 | 1,200 | 0.05% |
| Add to basket | 540,000 | 540 | 0.10% |
| Place order | 60,000 | 2,460 | 4.10% |
| All operations | 3,000,000 | 4,200 | 0.14% |
The headline figure of 0.14% looks tolerable, yet 4.1% of orders fail. If you treat the order call as the thing customers pay for, it deserves its own objective. Many teams define a handful of service level indicators, one for each journey that matters, instead of one number for the whole API.
A handful of simple habits keep the number trustworthy: publish the definition beside the figure, keep raw counts rather than only percentages, and review the numerator and denominator separately when the rate jumps. A rate can double because failures rose or because traffic fell, and the fix is different in each case.
Picking the window for the budget
A 30-day rolling window is the usual choice because it is long enough to smooth out a single bad hour and short enough that a regression is still visible. A calendar-month window resets cleanly, which is easier to explain to customers, but it lets a bad last day hide behind a fresh start the following morning.
Short windows are for paging; long windows are for planning. A common pattern is to page the on-call engineer when the burn rate over the last hour is very high and a confirming signal appears over the last few minutes, and to open a ticket when a slower burn lasts for a day or three. The budget itself stays on the 30-day figure.
Whatever you pick, write the window down next to the objective. A target such as 99.9% without a period is not something anyone can verify.
Where error rates mislead
- Low traffic makes rates jumpy. Two failures out of ten requests is 20%, which says little; use a request count threshold before alerting.
- Retries hide failures. If clients retry quietly, the user may never see an error, but your backend carries the extra load and the metric looks good.
- Averages across endpoints hide a broken one. A checkout call failing 30% of the time disappears when it is one percent of all traffic.
- Failures that never reach you are not counted. If the load balancer or the network drops a request before the application sees it, only client-side measurement catches it.
Common questions
How do you calculate API error rate?
Divide the number of failed requests by the total number of requests in the same period, then multiply by 100. For 1,500 failures out of 600,000 requests, the rate is 1,500 ÷ 600,000 × 100 = 0.25%.
What is an acceptable API error rate?
It depends on the service and the promise made to users. Many public APIs target 99.9% success, which is an error rate of 0.1% or less. Payment or safety systems aim lower, while internal tools often accept 1% or more.
What is an error budget in SRE?
An error budget is the amount of unreliability your SLO tolerates over a window. For a 99.9% objective across 10 million requests, the budget is 10,000 failures. Spending it too fast triggers a shift from features to reliability work.
Should 4xx responses count in the error rate?
Usually not for the service's own SLO, since 400, 401 or 404 normally come from the caller. Count 5xx responses and timeouts, and watch 4xx separately because a sudden rise can point to a broken release on the client side or a bad deploy that changed a contract.
How do you measure error rate in GraphQL?
Because most responses are HTTP 200, count operations whose response contains a non-empty errors array, and optionally those that returned partial data. Divide by total operations. Also track transport failures and timeouts, which never produce a GraphQL body.
Was this guide helpful?
Continue reading
View all blogsDecimal to Binary: Divide by 2, Read Backward
Convert any whole number to binary by halving it and reading the remainders from the bottom up. 45 becomes 101101; here is the method and a check.
5 min read
Download Time from File Size and Speed
Divide the file size in megabits by the speed in Mbps: 500 MB at 50 Mbps takes 80 seconds, because a byte is 8 bits and Mbps is not MB/s.
6 min read
Hex to Decimal: Convert Hexadecimal by Hand
To turn hex into decimal, multiply each digit by a power of 16 and add: 1A3F becomes 6,719. Includes the A to F table, colour codes and a quick shortcut.
5 min read




