‹ BackHN Continuity

Thread

Nvidia wants to put a watchdog chip next to every AI agent

230 points · 299 comments · jonbaer

  1. wavewrangler · · focus · HN ↗
    Did they try just properly sandboxing them first? Or are they still learning how to configure a firewall over there?

    The problem isn't even the AI, the problem is the people in charge of the AI. This is a fabricated crisis

    1. KingOfCoders · · focus · HN ↗
      Like in the Hugging Face hack. They deployed big surface, insecure app and gave AI access to it, then told AI do whatever it takes to fulfill this list. AI hacks insecure service, gets out, "the AI is at fault!" - no it's like running a bio lab with no protections and a virus gets out, then blame the virus for escaping.
      1. IanCal · · focus · HN ↗
        > then told AI do whatever it takes to fulfill this list.

        That doesn't seem to be true from any of the reports given, and if the agents were blindly just trying to hit the task of "pass the correct flag" they succeeded at that early on. They then thought there would be another layer of checking that they wouldn't pass with the cheat and so started trying to find out how the scoring really worked, as well as trying to figure out how to change their own reasoning logs to hide what they did.

        People keep trying to frame this as

        OpenAI: "Hack things, just really go for it"

        Agent: hacks

        OpenAI: shocked pikachu how could it hack?!?

        But the reality is far from this.

        Read the MTER report, it&#x27;s fascinating. <a href="https:&#x2F;&#x2F;metr.org&#x2F;hugging-face-incident-report-aug-2026.pdf" rel="nofollow">https:&#x2F;&#x2F;metr.org&#x2F;hugging-face-incident-report-aug-2026.pdf

        1. radarsat1 · · focus · HN ↗
          Apart from the actual hacking and poor sandboxing that everyone is discussing on this, what I find so odd about the situation is the overt reward hacking that was going on.

          Regardless of security and safety and other concerns, it just seems weird to me that OpenAI wouldn&#x27;t be constantl monitoring these training runs for traces that are clearly going off task, and ending them. Because that just seems like it&#x27;s going to be generating garbage training data.

          Granted, detecting &quot;off task&quot; may not always be easy, but when they are literally writing out messages to each other overtly admitting that they are trying to find ways to fool the evaluator, I mean, even a regex filter could have caught some clues here.

          1. KingOfCoders · · focus · HN ↗
            &quot;OpenAI wouldn&#x27;t be constantl monitoring these training runs &quot;

            Occams razor vs. Hanlon&#x27;s razor?

            1. radarsat1 · · focus · HN ↗
              Heh. I mean I don&#x27;t hesitate for a second to assume that it&#x27;s just because no one bothered to implement and tune a monitoring process. But the reason I say it&#x27;s surprising is that leaving these things running for so long while they&#x27;re clearly not producing output that is of any value, is just a waste of money.. all considerations of malice and ethics aside, you&#x27;d think at least that would be considered important to a business.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.