Retries, exponential back-off and jitter
Networks are unreliable by nature. Links drop, buffers fill, servers run hot. When you connect to a service you are betting that all the intermediate routers, switches, load balancers and firewalls will do their job. Most of the time they do, sometimes they do not. That is why retries exist.
Retries keep things alive when packets disappear or services return temporary errors. But retries done wrong create storms that make the outage worse.
Naïve retries
The most natural reaction to failure is to try again. That is what almost every client library does out of the box: if it fails, try again. It feels intuitive, harmless even. But in distributed systems the details of “try again” matter more than the attempt itself.
How many times do you retry? Some clients will keep looping forever, as if persistence alone will make the server come back to life. That is a recipe for runaway processes and wasted resources. A single misconfigured client like that can sit hammering a dead service until someone pulls the plug.
When do you retry? If the answer is “immediately,” you have just multiplied the traffic the failing system has to deal with. Imagine a busy API that starts returning 503 because it is overloaded. Every client retries instantly, adding more requests on top of the ones that already pushed it over the edge. What could have been a short-lived blip now drags into a full outage because the retries themselves keep the system pinned down.
Under what conditions do you retry? Not all errors are equal. A timeout probably means the network dropped the packet or the server was too slow, worth trying again. A 500 error is usually a transient server-side problem, also worth retrying. But a 404 will never succeed no matter how many times you send it. Authentication failures will not magically fix themselves. Retrying those just eats cycles and makes logs noisy.
What if the server is just slow? This is the scenario that trips up most naïve retry implementations. If responses are delayed and the client does not wait long enough before deciding the request failed, it will send another one. Now the server has to process two identical requests instead of one, on top of being slow in the first place. Multiply that across thousands of clients and the retries themselves become the dominant source of load. You do not just fail to fix the problem, you make it worse.
And then there is cost. In the cloud every request has a price tag. Every retry burns CPU, bandwidth, database connections, TLS handshakes. On small systems that might be invisible, but at scale it shows up as a sharp increase in cloud bills. Teams have been surprised by five-figure invoices because of badly tuned retries.
The pattern is clear: naïve retries do not just fail to help, they actively make systems less reliable. They turn transient problems into outages, overload services that could have recovered on their own, and hide bugs behind layers of noise. The instinct to “just try again” is correct, but without limits, pacing and context it becomes dangerous.
What can we do?
If naïve retries make things worse, what does it look like to get them right? The answer is a mix of pacing, randomness, and awareness of failure modes.
The first piece is exponential back-off. Instead of retrying immediately or at a fixed interval, the client waits progressively longer between attempts. The first retry might happen after 200 milliseconds, the next after 400, the next after 800, and so on. This simple pattern gives the remote service time to recover. A database under load can clear its backlog, a congested link can drain its queue, a container can restart. Without that breathing room, the service is pinned down by waves of retries that arrive faster than it can heal.
The second piece is jitter. Exponential back-off without jitter means every client is still in sync. If ten thousand clients all fail at the same moment they will all retry after 200ms, then all at 400ms, then all at 800ms. That is the thundering herd problem. The way out is to randomise the wait time. Instead of always waiting exactly 800ms, each client picks a random delay up to 800ms. Some retry earlier, some later, and the retries are spread more evenly across time. The server sees a smoother trickle of traffic instead of spikes that keep knocking it over.
There are variations of jitter. Full jitter chooses a random value anywhere between zero and the back-off delay. Equal jitter centres the random value around half the back-off. Decorrelated jitter uses a more complex calculation to avoid sudden jumps and smooth the pattern further. Large-scale platforms like AWS use decorrelated jitter in their SDKs because it works well when millions of clients may retry at once.
Exponential back-off and jitter form the backbone, but other practices matter as well. Always limit the number of retries. Three to five attempts is common, sometimes fewer if latency is critical. Without a cap, clients can loop forever and create unnecessary load. Always set a maximum delay. If your back-off keeps doubling unchecked, you may end up with clients waiting for minutes before giving up. That is wasted time for the user and often pointless for the system.
Choose retry conditions carefully. Retry on network timeouts, TCP resets, or 5xx errors from the server. Do not retry on permanent errors like 404, 401, or invalid request formats. Retrying those will never change the outcome. Knowing which errors are transient and which are permanent is critical to tuning the retry policy.
Add visibility. Retries should be measured and logged. Spikes in retry counts are often the earliest warning sign of a larger outage. If you only look at the success rate you may miss the fact that many of those successes are second or third attempts. A system that “works” but needs retries constantly is not healthy.
Finally, consider when not to retry at all. Circuit breakers are a common pattern for this. If a service is known to be down, stop sending requests until it shows signs of life again. This prevents floods of useless traffic and allows upstream systems to stabilise. Bulkheads and rate limits are also relevant, making sure that one misbehaving dependency does not drag down the rest of the platform.
Good retry behaviour is about restraint. Slow down instead of rushing, spread out instead of synchronising, stop when the problem is permanent, and pay attention to the signals retries give you. Systems designed with these practices do not just survive failures, they degrade gracefully under pressure instead of collapsing.
When to stop retrying
Retries are not free. In distributed systems every request usually holds some resource. It might be a database transaction, a lock in a coordination service, a file handle, or a slot in a thread pool. If the client keeps retrying while the resource stays allocated, the retries prevent recovery instead of helping it.
Knowing when to stop retrying is just as important as knowing how to retry. A transaction that fails after holding a lock should not be retried endlessly by the same process, because the lock itself may still be in place. An RPC call that times out after opening a session may leave the session dangling on the server. Retrying too many times can create a trail of half-finished work that makes the real outage worse.
The safe pattern is to bound retries not just by count and delay, but also by context. Stop retrying if the operation is holding a scarce resource. Stop retrying if the request is part of a larger workflow that has already timed out. Stop retrying if the dependency is known to be down for a long time and upstream systems should fail fast instead of waiting.
Retries should serve the system, not the other way around. Once they start keeping resources occupied that could be freed for others, it is time to stop.
Monitoring
The effects of retries are easy to miss if you only look at headline metrics. A service that starts slowing down will show longer response times. If clients retry aggressively those retries turn into extra requests. In monitoring this appears as a sudden increase in requests per second at the same time as latency spikes.
It is tempting to read that as higher load causing slower responses. In reality the causality is reversed: the service slowed down first, and the retries created the load spike. The danger is that you focus on scaling up capacity when the real issue is retry behaviour. The two signals, slow responses and increased RPS, reinforce each other and make the problem look like a traffic surge rather than a retry storm.
Clear instrumentation helps untangle this. Track retry counts separately from raw request counts. Compare initial requests with total requests including retries. Watch for cases where a dip in success rate is followed by a jump in overall RPS. Those patterns are the signature of compounded retries turning a partial fault into a system-wide incident.
Implications
At small scale retries are invisible. At cloud scale retries define failure behaviour. Many incidents are not caused by the first fault but by the retries that followed. Clients overwhelm the very systems they are trying to reach. APIs collapse under their own load. Bills climb as wasted calls pile up. Dependencies fail in a chain as every service retries the one below it.
Handled carefully, retries make systems resilient. Handled carelessly, they create denial of service against your own infrastructure. The difference comes down to pacing, randomness, sensible limits, knowing when to stop, and awareness of which errors are worth retrying in the first place.
Member discussion