‹ BackHN Continuity

Thread

How Uber Protects Against Retry Storms

120 points · 49 comments · iscmt

  1. aftbit · · focus · HN ↗
    I'd be interested to hear other strategies in this space. I've done the naive thing of allowing retries everywhere, and gotten into retry storms. When I was next presented with the problem, I tried the other naive thing of only allowing retries from the very top level service, which led me to redoing absolutely tons of work for each failure. What's a nice middle path that doesn't add too much complexity?
    1. CBLT · · focus · HN ↗
      There's a good amount of literature about this (check the other comments), but you can vastly simplify this into two things you need to do:

      1. Your service that retries should have some retry budget. This is a good place to be "smart", because you can reason entirely locally instead of turning it into a distributed systems problem. The best library I've seen for this was doing Exponential Moving Average of requests per second sent down that pipe (not counting retries) and only allowing 20% more requests per second as retries, total. Each individual request could be retried 3 times. This was critical as it bounds the additional load from retries.

      2. Whenever a service retries but has to give up, the error it sends to its callers should never be retried. There has to be some agreement that that HTTP code will never be retried. This prevents the multiplicative factor of retry on top of retry, which is why those storms can generate so much load.

      Everything else is nice-to-have, but those two alone should bound the total requests you get in a retry storm.

      1. sandeepkd · · focus · HN ↗
        I get a feeling of dejavu for this one. Most of my experience has been in Java and in most places I worked in the past we had this hierarchy of exception classification which gets reflected into the http status codes as well. On high level the HTTP status codes in case of errors are already classified as re-tryable or not, the convention varies globally but can be adopted in a standard manner within a company.

        The reason why I brought up the exception propagation is cause within a large enough service with multiple layers of depth the exception hierarchy provides with similar context.

        The hard part is not implementing something like this, its about maintaining it consistently across every new change. With small product teams this architecture concept/convention/constraint can easily get lost/forgotten and what you are left with is a theoretical system which does not works as desired when the storm comes

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.