Engineering
7 min read
Why I stopped writing my own retry loop
Every project I have worked on has a retry loop somebody wrote in a hurry. Here is what a good one looks like, and why you should stop writing yours.

The first retry loop I ever wrote was four lines long. It waited a second, tried again, and gave up after three attempts. It worked for two years and then took down a payments service for forty minutes.
What goes wrong
When a server is struggling, every client retrying at the same moment makes it worse. A fixed delay means all of your clients come back together. The fix is old and simple: wait longer each time, and add a random amount so nobody is in step.
Three rules I now follow
Retry only what is safe to repeat. Put a ceiling on total time, not just on the number of attempts. And log every retry with the reason, because the day you need it you will not be able to reproduce the failure.
Then stop writing it
None of this is interesting, and that is the point. A retry policy should be a setting you can read, not a function you have to trust. That is why Plumbline ships one and nothing else.
Written by
Iri Halvorsen, software engineer in Bergen
More / ls -t ./writing

