‹ BackHN Continuity

Thread

How Uber Protects Against Retry Storms

120 points · 49 comments · iscmt

  1. aftbit · · focus · HN ↗
    I'd be interested to hear other strategies in this space. I've done the naive thing of allowing retries everywhere, and gotten into retry storms. When I was next presented with the problem, I tried the other naive thing of only allowing retries from the very top level service, which led me to redoing absolutely tons of work for each failure. What's a nice middle path that doesn't add too much complexity?
    1. jeffbee · · focus · HN ↗
      Just limiting your retry budget to 1% of normal rates using a client-local token bucket with no distributed coordination will eliminate the possibility of long-lived retry storms.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.