‹ BackHN Continuity

Thread

Livenerf: Has Opus 5.5 been nerfed yet?

922 points · 392 comments · bryan0

  1. sheepscreek · · focus · HN ↗
    > It could also mean nothing happened and people are pattern-matching on noise.

    It is definitely not this. Anthropic has thousands of employees making probably > 100,000+ tiny changes across the entire stack and infrastructure everyday.

    The compounding effect can definitely cause temporary regressions in some domain or use-case that doesn’t have really good coverage in their internal tests.

    What makes this particularly challenging in the case of LLMs is how changing some language in the prompt can vastly affect the output.

    But this is less true as models become larger and additional parameters are able to capture each and every possible nuance of the language. I’d say caching and cache tuning or token optimization/tweaks to thinking are the biggest culprits today.

    1. zahlman · · focus · HN ↗
      > It is definitely not this. Anthropic has thousands of employees making probably > 100,000+ tiny changes across the entire stack and infrastructure everyday.

      I don't follow. That sounds to me exactly like a reason why it could be people pattern-matching on noise: because there is a lot of noise in which a matchable pattern could emerge.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.