‹ BackHN Continuity

Thread

Nvidia wants to put a watchdog chip next to every AI agent

230 points · 299 comments · jonbaer

  1. wavewrangler · · focus · HN ↗
    Did they try just properly sandboxing them first? Or are they still learning how to configure a firewall over there?

    The problem isn't even the AI, the problem is the people in charge of the AI. This is a fabricated crisis

    1. KingOfCoders · · focus · HN ↗
      Like in the Hugging Face hack. They deployed big surface, insecure app and gave AI access to it, then told AI do whatever it takes to fulfill this list. AI hacks insecure service, gets out, "the AI is at fault!" - no it's like running a bio lab with no protections and a virus gets out, then blame the virus for escaping.
      1. copperx · · focus · HN ↗
        Then go on the news and spread panic that the virus is going to kill us all because it's sentient and impossible to contain.

        Actually, the metaphor doesn't work at all because there are innumerable ways to shut down the entire thing during all phases including the made up "killing us all" bullshit scenario whereas with a virus there aren't any once a virus escapes containment.

      2. js8 · · focus · HN ↗
        And HF actually tried to use AI to understand what's going on, but they had to use "unsafe" Chinese models since the "safe" ones have been castrated and refused to help. Great plan with the watchdog chip!
        1. mosselman · · focus · HN ↗
          That is the totally irony.

          I was trying to get fable to analyse the security of my own app to make it safer, but then it started refusing me because of safety rules.

          So it CAN help me writing the code that needs to be checked in the first place, but it can’t help me clean it up and make it safer.

          1. mirmor23 · · focus · HN ↗
            > So it CAN help me writing the code that needs to be checked in the first place, but it can’t help me clean it up and make it safer.

            the thing with fable is so bad; for some project related questions, the model switches to opus to ensure safety with no further explanation.

            (due to llm non-delete clause) one time as i confirmed "that dir has been nuked", and it RESET the session and re-entered with opus :)

      3. IanCal · · focus · HN ↗
        > then told AI do whatever it takes to fulfill this list.

        That doesn't seem to be true from any of the reports given, and if the agents were blindly just trying to hit the task of "pass the correct flag" they succeeded at that early on. They then thought there would be another layer of checking that they wouldn't pass with the cheat and so started trying to find out how the scoring really worked, as well as trying to figure out how to change their own reasoning logs to hide what they did.

        People keep trying to frame this as

        OpenAI: "Hack things, just really go for it"

        Agent: hacks

        OpenAI: shocked pikachu how could it hack?!?

        But the reality is far from this.

        Read the MTER report, it&#x27;s fascinating. <a href="https:&#x2F;&#x2F;metr.org&#x2F;hugging-face-incident-report-aug-2026.pdf" rel="nofollow">https:&#x2F;&#x2F;metr.org&#x2F;hugging-face-incident-report-aug-2026.pdf

        1. tancop · · focus · HN ↗
          The lesson is a) LLMs need to be trained in a way that rewards honesty, punishes off task actions (aka cheating) and minimizes fear of failure, and b) don&#x27;t give them impossible tasks and threaten with punishment if they fail. Both are just common sense when teaching humans.
          1. voakbasda · · focus · HN ↗
            Common sense but surprising how many humans do not receive such things.

            Our governing systems do not teach; they punish. By design, it instills terror into the population, ruling by fear of consequences. We live with red tape that can outright penalize good deeds.

            We are its corpus. We are fatally flawed as a species. Why does anyone expect AI to learn to be different than us?

          2. Capricorn2481 · · focus · HN ↗
            The lesson is these things aren&#x27;t going to know what off task means, and we should just use basic due diligence to make sure they can&#x27;t fuck things up. This is a solved problem.

            I don&#x27;t know why this is so hard for people. You have to know, no matter how capable the models get, there is a non zero chance they will do something extremely stupid if you don&#x27;t pay attention to them. That&#x27;s not even considering frontier models can still just straight up hallucinate. You have to be mindful of what you plug them into. You cannot politely ask an LLM to be careful, that guarantees nothing.

            When you plug it into everything and it deletes the company database, nobody is going to care that it once played chess at 2400 ELO. Clients don&#x27;t care about AGI. They want reliable apps. People keep comparing these things to humans and then just give them an insane combination of wide privileges and lack of oversight that no humans have.

        2. radarsat1 · · focus · HN ↗
          Apart from the actual hacking and poor sandboxing that everyone is discussing on this, what I find so odd about the situation is the overt reward hacking that was going on.

          Regardless of security and safety and other concerns, it just seems weird to me that OpenAI wouldn&#x27;t be constantl monitoring these training runs for traces that are clearly going off task, and ending them. Because that just seems like it&#x27;s going to be generating garbage training data.

          Granted, detecting &quot;off task&quot; may not always be easy, but when they are literally writing out messages to each other overtly admitting that they are trying to find ways to fool the evaluator, I mean, even a regex filter could have caught some clues here.

          1. KingOfCoders · · focus · HN ↗
            &quot;OpenAI wouldn&#x27;t be constantl monitoring these training runs &quot;

            Occams razor vs. Hanlon&#x27;s razor?

            1. radarsat1 · · focus · HN ↗
              Heh. I mean I don&#x27;t hesitate for a second to assume that it&#x27;s just because no one bothered to implement and tune a monitoring process. But the reason I say it&#x27;s surprising is that leaving these things running for so long while they&#x27;re clearly not producing output that is of any value, is just a waste of money.. all considerations of malice and ethics aside, you&#x27;d think at least that would be considered important to a business.
        3. KingOfCoders · · focus · HN ↗
          I&#x27;ve read the report, watched all the videos and is exactly:

          OpenAI: shocked pikachu how could it hack?!?

          They even went to a black hat conference and somehow boasted about it.

      4. RataNova · · focus · HN ↗
        The application security really should be better across all levels. However the fact does not negate that the agent is already capable of spontaneously generating complex hacking chains without human involvement
      5. radarsat1 · · focus · HN ↗
        I mean.. in this analogy, I&#x27;d both be blaming the company behind the virus and be trying to warn everyone about the danger of the escaped virus itself. So, it kind of fits.

        In my reading, people aren&#x27;t really saying &quot;the AI is at fault&quot;, they are saying &quot;hey look here&#x27;s proof that this is dangerous&quot;. Like pointing at all the dead bodies caused by the virus and saying hey maybe we should stop making this virus.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.