I'd be interested to hear other strategies in this space. I've done the naive thing of allowing retries everywhere, and gotten into retry storms. When I was next presented with the problem, I tried the other naive thing of only allowing retries from the very top level service, which led me to redoing absolutely tons of work for each failure. What's a nice middle path that doesn't add too much complexity?
aftbit · · focus · HN ↗
otterley · · focus · HN ↗
<a href="https://builder.aws.com/content/3EumjoZascWd1oZiEgL8ORlv3qE/timeouts-retries-and-backoff-with-jitter" rel="nofollow">https://builder.aws.com/content/3EumjoZascWd1oZiEgL8ORlv3qE/... (originally published 2020, republished 2026)
<a href="https://docs.aws.amazon.com/sdkref/latest/guide/feature-retry-behavior.html#retry-quota-token-bucket" rel="nofollow">https://docs.aws.amazon.com/sdkref/latest/guide/feature-retr...
<a href="https://aws.amazon.com/blogs/developer/announcing-updated-retry-behavior-for-aws-sdks-and-tools/" rel="nofollow">https://aws.amazon.com/blogs/developer/announcing-updated-re... (2026)