Cache hit rate:
what it is and how it changes response time
See why the last few percentage points of hit rate matter far more than the headline figure suggests.
Calcylator Editorial Team
Updated · 5 min read
Hit rate and miss rate defined
A cache request is either a hit, answered from the stored copy, or a miss, which has to go to the slower origin such as a database, disk or upstream server. Hit rate is the fraction of requests that are hits, and miss rate is the remainder.
- hits:
- requests served from the cache
- misses:
- requests that had to go to the origin
Hits
9,200
Misses
800
Hit rate
92% (miss rate 8%)
9,200 ÷ (9,200 + 800) = 9,200 ÷ 10,000 = 0.92.
Count the same thing in both terms. CDNs sometimes report byte hit rate (share of bytes served from cache) alongside request hit rate, and the two can differ a lot when large files behave unlike small ones.
What hit rate does to average latency
- h:
- hit rate as a decimal
- t_cache:
- time to answer from the cache
- t_origin:
- time for a miss, including the trip to the origin
Take a cache that answers in 2 ms and an origin that takes 120 ms. The average over a mix of requests moves as follows.
| Hit rate | Average latency | Origin load at 1,000 req/s |
|---|---|---|
| 80% | 25.6 ms | 200 req/s |
| 92% | 11.4 ms | 80 req/s |
| 96% | 6.7 ms | 40 req/s |
| 99% | 3.2 ms | 10 req/s |
Going from 80 to 92 percent more than halves latency, and the last stretch from 96 to 99 percent halves it again. The gain looks small in percentage points but is large because the origin cost is spread over a shrinking number of requests.
Misses decide how hard the origin works
The better lens is the miss rate, since that is the traffic the database or backend still has to handle. Moving from 92 to 96 percent hits cuts the miss rate from 8 to 4 percent, which halves origin load. The same two percentage points near 50 percent would make little difference to it.
This is why teams chase the tail. A backend sized for 80 requests per second has far more room than one sized for 200, and a surge in traffic is absorbed by the cache, not the database.
Hit rates compound across layers
Most requests pass through more than one cache: the browser, a CDN, then an application cache in front of the database. What reaches the origin is the product of the miss rates at each layer, so modest hit rates multiply into a large reduction.
Browser cache hit rate
40%
CDN hit rate on what reaches it
90%
Page loads
10,000
Requests reaching the origin
600 (6%)
Browser misses 60% × CDN misses 10% = 0.06, and 0.06 × 10,000 = 600.
Remember that each layer has its own definition and its own reporting. A CDN's hit ratio only covers the traffic it sees, so the overall figure has to be built from the layers, not read from one dashboard.
Common reasons for a low hit rate
- Cache keys that include noise such as tracking parameters, timestamps or user IDs, so identical content gets many keys.
- Time-to-live set too short for content that rarely changes.
- Cache too small, so useful entries are evicted before they are reused.
- Personalised or authenticated responses marked as uncacheable when parts of them could be shared.
- Very long-tail traffic where most items are requested once.
Fixing key normalisation usually gives the biggest jump for the least effort. Check what makes up the key before you buy more memory.
When the number can mislead
A high hit rate on trivial requests can hide a poor result on expensive ones. If 95 percent of hits are for tiny static files while the costly API responses almost always miss, the origin is still struggling.
Segment the metric by route, content type or customer tier. Also note what counts as a miss: stale-while-revalidate and conditional requests that get a 304 are handled differently by different systems, so confirm the definition before you compare tools or set targets.
Finally, the best hit rate is not always 100 percent. Content that must be fresh, such as a balance or inventory count, may deliberately bypass the cache.
Matching symptoms to causes
When the rate disappoints, the pattern tells you where to look before you change any settings.
| Symptom | Likely cause | First thing to try |
|---|---|---|
| Low rate right after a deploy | Cold cache or changed keys | Warm popular keys; keep keys stable between versions |
| Rate drops at peak traffic | Cache too small, entries evicted early | Raise memory or shorten the stored value size |
| Many near-duplicate keys | Unneeded query parameters in the key | Normalise and sort parameters, strip tracking ones |
| Rate falls on a schedule | TTLs expire together | Add jitter to TTLs so expiry is spread out |
Eviction policy matters too. Least-recently-used suits most workloads, but a one-time scan of the whole catalogue can push out useful entries, which is why some systems protect a hot set or admit new items more selectively.
The cost view and what to alert on
Hit rate also translates into money. Suppose a site serves 10 TB a month. With a byte hit rate of 92 percent, 0.8 TB is pulled from the origin; at 96 percent, only 0.4 TB. If egress or compute behind the origin is charged by volume, the second figure halves that bill.
For operations, a steady number matters more than a perfect one. Alert on a sudden fall relative to the usual level for that hour, since a drop from 94 to 85 percent usually signals a bad deploy, a changed key or a purge, long before users complain. Track the metric per route or content class so that a healthy aggregate cannot hide one struggling endpoint.
- Record hits, misses and bypasses separately; a bypass is not a miss in the same sense.
- Compare against the same weekday and hour, not just the previous minute.
- Pair hit rate with origin latency and error rate on the same dashboard.
One last caution concerns freshness. Raising the lifetime of cached entries is the cheapest way to improve the hit rate, but it also lengthens the time that users can see out-of-date content. Decide per content type how stale is acceptable, use explicit invalidation for the few items that must change at once, and treat the hit rate as one half of a pair, alongside a measure of how often users actually received something older than they should have.
Common questions
How do you calculate cache hit rate?
Divide hits by total requests, which is hits plus misses. For 9,200 hits and 800 misses, that is 9,200 ÷ 10,000 = 92 percent. The miss rate is the remaining 8 percent.
What is a good cache hit rate?
It depends on the workload. Static assets on a CDN often exceed 90 percent, while personalised API data may sit much lower. Judge by what the miss traffic costs your origin and users, not by one universal target.
How does cache hit rate affect latency?
Average latency is hit rate times cache time plus miss rate times origin time. With 2 ms and 120 ms, 92 percent gives 11.4 ms and 99 percent gives 3.2 ms, so the final points matter most.
What is the difference between hit ratio and byte hit ratio?
Hit ratio counts requests; byte hit ratio counts data volume. A cache can hit most small files yet miss large ones, giving a high request ratio but a low byte ratio. Bandwidth costs follow the byte figure.
How can I increase my cache hit rate?
Normalise cache keys, remove irrelevant query parameters, lengthen the TTL where content allows, increase cache size, and warm the cache before traffic. Review misses by route to find the few that dominate.
Was this guide helpful?
Continue reading
View all blogsIP Address Subnet Calculation: Network, Broadcast, Hosts
Give an IPv4 address and prefix and you can derive its network, broadcast and host range. 10.20.77.130/27 sits in 10.20.77.128 to .159 with 30 hosts.
4 min read
Subnet Mask From a Prefix: How /24 Becomes 255.255.255.0
A /24 prefix means 24 leading one-bits, which is 255.255.255.0 and 254 usable hosts. Learn to convert any prefix and find the network and broadcast address.
5 min read
Download Time from File Size and Speed
Divide the file size in megabits by the speed in Mbps: 500 MB at 50 Mbps takes 80 seconds, because a byte is 8 bits and Mbps is not MB/s.
6 min read




