I'm suspicious of load shedding not mentioned in the article. Combine that with exp backoff in the caller and you got yourself a pretty robust starting point
429 (and sometimes 503) errors returned by servers might well be a symptom of intentional load shedding. Perhaps it's just not explicitly called out as a server behavior that induces client retries.
I’d like to mention it since Uber has a really cool load shedder [1], also implemented similarly by Netflix [2] and failsafe-go [3]. It basically looks for points where more requests per second suddenly cause a significant increase in latency and calls that the concurrency limit.
maxchisto · · focus · HN ↗
otterley · · focus · HN ↗
hxtk · · focus · HN ↗
1: <a href="https://www.uber.com/us/en/blog/cinnamon-using-century-old-tech-to-build-a-mean-load-shedder/" rel="nofollow">https://www.uber.com/us/en/blog/cinnamon-using-century-old-t...
2: <a href="https://github.com/Netflix/concurrency-limits" rel="nofollow">https://github.com/Netflix/concurrency-limits
3: <a href="https://failsafe-go.dev/adaptive-limiter/" rel="nofollow">https://failsafe-go.dev/adaptive-limiter/