‹ BackHN Continuity

Thread

How Uber Protects Against Retry Storms

120 points · 49 comments · iscmt

  1. aftbit · · focus · HN ↗
    I'd be interested to hear other strategies in this space. I've done the naive thing of allowing retries everywhere, and gotten into retry storms. When I was next presented with the problem, I tried the other naive thing of only allowing retries from the very top level service, which led me to redoing absolutely tons of work for each failure. What's a nice middle path that doesn't add too much complexity?
    1. tregoning · · focus · HN ↗
      <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Exponential_backoff" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Exponential_backoff
      1. sroussey · · focus · HN ↗
        Yes, exponential back off and jitter are the first things to work on, and good if you don’t have a better signal (like loss of network).

        Also, a simple signal status server or queue system helps to keep global state such that everyone doesn’t retry all at once.

        If you have a central error rate server you can skip your retry based on the error rate (100% error rate, don’t retry, etc).

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.