I'd be interested to hear other strategies in this space. I've done the naive thing of allowing retries everywhere, and gotten into retry storms. When I was next presented with the problem, I tried the other naive thing of only allowing retries from the very top level service, which led me to redoing absolutely tons of work for each failure. What's a nice middle path that doesn't add too much complexity?
Just limiting your retry budget to 1% of normal rates using a client-local token bucket with no distributed coordination will eliminate the possibility of long-lived retry storms.
aftbit · · focus · HN ↗
jeffbee · · focus · HN ↗