I'd be interested to hear other strategies in this space. I've done the naive thing of allowing retries everywhere, and gotten into retry storms. When I was next presented with the problem, I tried the other naive thing of only allowing retries from the very top level service, which led me to redoing absolutely tons of work for each failure. What's a nice middle path that doesn't add too much complexity?
aftbit · · focus · HN ↗
tregoning · · focus · HN ↗
anonymars · · focus · HN ↗
mitxela · · focus · HN ↗
Microsoft breaks all Old New Thing links every few years so it's necessary to post the title so the right post can still be found.