‹ BackHN Continuity

Thread

Livenerf: Has Opus 5.5 been nerfed yet?

922 points · 392 comments · bryan0

  1. jug · · focus · HN ↗
    We also have Nerf Bench:

    <a href="https:&#x2F;&#x2F;www.bridgebench.ai&#x2F;nerf-bench" rel="nofollow">https:&#x2F;&#x2F;www.bridgebench.ai&#x2F;nerf-bench

    They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They&#x27;re currently tracking Opus 5.5 and GPT-6 Astra.

    This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it&#x27;s often about honeymoon effects.

    1. rplnt · · focus · HN ↗
      &gt; I personally think people sense nerfs more often than they happen and that it&#x27;s often about honeymoon effects.

      I believe in temporary nerfs. Operators reducing quality significantly to increase throughput for whatever reason (high demand?). Same session, model being completly incapable, and it being back to normal the next day. Experienced that with Anthropic models way too many times. Never on weekends, usually during US work days.

      It&#x27;s been a few months since I last recall this though, must have been pre-opus-5. And I know there are benchmarks for this as well, hence &quot;believe&quot;.

      1. transcriptase · · focus · HN ↗
        I vividly remember when ChatGPT3.5 went fully mainstream, there were times where within minutes you would realize they were only serving up idiot mode and there was no point trying to do much until demand died down and they swapped back to the non-quantized version.

        People called it lazy mode, in that instead of writing the script you asked for it would basically tell you to learn to code then check out xyz topics to tackle the problem.

        1. katzenq · · focus · HN ↗
          It should have stayed that way. Be a good search engine and encourage the human to do the work themselves.
          1. AbsurdCensor · · focus · HN ↗
            Wouldn&#x27;t that be like a calculator saying &#x27;pull out a math book&#x27; when you try running calculations?
            1. katzenq · · focus · HN ↗
              It&#x27;s like doing the calculations yourself (with whatever tools you want) and knowing what you&#x27;re doing, steering the process yourself, rather than begging someone to give you the result so you can then show it off like you did the work. I can understand and explain what my calculator is doing and I&#x27;m not offloading decisions to it. Using a language model to shit out projects that you can&#x27;t fully understand the structure of yourself is insane. Ceding control of any design decisions to a language model is insane. They&#x27;re fine as second-order autocomplete and semantic search engines. Humans should be handling design and implementation entirely, making things for other humans. An LLM running in a loop can&#x27;t build humane systems.
              1. sejje · · focus · HN ↗
                &gt; Humans should be handling design and implementation entirely, making things for other humans. An LLM running in a loop can&#x27;t build humane systems.

                I tell the LLM what to do in its loop. It builds it. I tweak it until it is perfect. I care a lot about UI.

                I&#x27;m mostly building my own UIs lately, but I find no problem with the development loop. I&#x27;m building much better UIs, because it&#x27;s way easier to test things, and scrap things that I thought would work, but don&#x27;t. It all happens in a matter of minutes.

                Humans should use more llms.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.