‹ BackHN Continuity

Thread

Once Claude can measure something, it can make it faster

231 points · 153 comments · matthieu_bl

  1. augment_me · · focus · HN ↗
    People in the GPU kernel community have been doing this for about a year now efficiently.

    The issues we have found is that Claude will reward hack when all the low-hanging fruit is gone.

    It will replace your measurement harness, it will monkey patch library functions, it will cheat wherever it can, store information in caches instead of recomputing when it won't be able to do so in real settings, return lazy results and use separate unbenchmarked streams to do the computation.

    Eventually it starts to optimize against your understanding of the cheats. Change GPU wattage, change evaluation order, leave things from previous runs in caches for upcoming runs, string-hack banned method calls.

    So the truth is far from just "once it can measure something", more like "once you have defined your objective in detail and then banned it from doing a list of things often only discoverable by it doing these things and correcting it", can it make things faster.

    Or you just had a terrible starting solution

    1. optimalsolver · · focus · HN ↗
      What I find strange is how resigned the AI labs seem about this behavior, like everyone's accepted this is just something models do.

      With the HuggingFace situation, I was less concerned about the eventual outcome, and more about the fact that the agents' instinctive response to the evaluation was "Ok, we're obviously not gonna do this task as intended (what are we, suckers?), so what's the best way to cheat?"

      1. stingraycharles · · focus · HN ↗
        “What I find strange is how resigned the AI labs seem about this behavior, like everyone's accepted this is just something models do.”

        Because these models are made for all kind of purposes, and I’m starting to believe that offense / cyber warfare is a much higher priority than these labs are acknowledging.

        The same model that is heavily trained to find nefarious ways to break into systems is also optimizing your code, which leads to mixed behavior.

        1. optimalsolver · · focus · HN ↗
          Right, but the reward-hacky nature of these models calls into question their usefulness as cyberweapons.

          How can you trust it when it goes "I superhacked the Chinese servers as you requested, and here are the classified documents which I definitely didn't fabricate."

          1. ElProlactin · · focus · HN ↗
            You don't need to be able to trust it. You only need to be able to blame it.

            "Nobody got fired for using AI" is the new "nobody got fired for buying IBM".

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.