‹ BackHN Continuity

Thread

How Uber Protects Against Retry Storms

120 points · 49 comments · iscmt

  1. aftbit · · focus · HN ↗
    I'd be interested to hear other strategies in this space. I've done the naive thing of allowing retries everywhere, and gotten into retry storms. When I was next presented with the problem, I tried the other naive thing of only allowing retries from the very top level service, which led me to redoing absolutely tons of work for each failure. What's a nice middle path that doesn't add too much complexity?
    1. sroussey · · focus · HN ↗
      So many variables, but the simple thing is to set things up like normal rate limiting (which you would want to do anyways). The one generating the errors passes back a retry time. You can add jitter here, tell low priority requests to wait longer, etc.

      BTW: do keep track of priority. It’s like having a database that gets flooded with connections and won’t allow new ones in—but will for admin users (btw, it did not used to be that way in the early days of MySQL).

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.