‹ BackHN Continuity

Thread

How Uber Protects Against Retry Storms

120 points · 49 comments · iscmt

  1. aftbit · · focus · HN ↗
    I'd be interested to hear other strategies in this space. I've done the naive thing of allowing retries everywhere, and gotten into retry storms. When I was next presented with the problem, I tried the other naive thing of only allowing retries from the very top level service, which led me to redoing absolutely tons of work for each failure. What's a nice middle path that doesn't add too much complexity?
    1. cyberax · · focus · HN ↗
      One good option that is not (yet?) mentioned here is a deadline for retries. You can cap the request duration by, say, 500ms and pass the remaining time budget to downstream services.

      This can be done via an HTTP header and enforced by the middleware.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.