‹ BackHN Continuity

Thread

Livenerf: Has Opus 5.5 been nerfed yet?

922 points · 392 comments · bryan0

  1. sheepscreek · · focus · HN ↗
    > It could also mean nothing happened and people are pattern-matching on noise.

    It is definitely not this. Anthropic has thousands of employees making probably > 100,000+ tiny changes across the entire stack and infrastructure everyday.

    The compounding effect can definitely cause temporary regressions in some domain or use-case that doesn’t have really good coverage in their internal tests.

    What makes this particularly challenging in the case of LLMs is how changing some language in the prompt can vastly affect the output.

    But this is less true as models become larger and additional parameters are able to capture each and every possible nuance of the language. I’d say caching and cache tuning or token optimization/tweaks to thinking are the biggest culprits today.

    1. jibalt · · focus · HN ↗
      > It is definitely not this.

      I don't think you understand what "this" is--or rather, the "thing" that didn't happen in "nothing happened".

      > Anthropic has thousands of employees making probably > 100,000+ tiny changes across the entire stack and infrastructure everyday.

      So, not the sort of thing referred to.

      > The compounding effect can definitely cause temporary regressions in some domain or use-case that doesn’t have really good coverage in their internal tests.

      Yes, but you claimed that this definitely did happen. But the whole point is to determine whether it did.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.