‹ BackHN Continuity

Thread

Nvidia wants to put a watchdog chip next to every AI agent

230 points · 299 comments · jonbaer

  1. wavewrangler · · focus · HN ↗
    Did they try just properly sandboxing them first? Or are they still learning how to configure a firewall over there?

    The problem isn't even the AI, the problem is the people in charge of the AI. This is a fabricated crisis

    1. KingOfCoders · · focus · HN ↗
      Like in the Hugging Face hack. They deployed big surface, insecure app and gave AI access to it, then told AI do whatever it takes to fulfill this list. AI hacks insecure service, gets out, "the AI is at fault!" - no it's like running a bio lab with no protections and a virus gets out, then blame the virus for escaping.
      1. IanCal · · focus · HN ↗
        > then told AI do whatever it takes to fulfill this list.

        That doesn't seem to be true from any of the reports given, and if the agents were blindly just trying to hit the task of "pass the correct flag" they succeeded at that early on. They then thought there would be another layer of checking that they wouldn't pass with the cheat and so started trying to find out how the scoring really worked, as well as trying to figure out how to change their own reasoning logs to hide what they did.

        People keep trying to frame this as

        OpenAI: "Hack things, just really go for it"

        Agent: hacks

        OpenAI: shocked pikachu how could it hack?!?

        But the reality is far from this.

        Read the MTER report, it&#x27;s fascinating. <a href="https:&#x2F;&#x2F;metr.org&#x2F;hugging-face-incident-report-aug-2026.pdf" rel="nofollow">https:&#x2F;&#x2F;metr.org&#x2F;hugging-face-incident-report-aug-2026.pdf

        1. tancop · · focus · HN ↗
          The lesson is a) LLMs need to be trained in a way that rewards honesty, punishes off task actions (aka cheating) and minimizes fear of failure, and b) don&#x27;t give them impossible tasks and threaten with punishment if they fail. Both are just common sense when teaching humans.
          1. voakbasda · · focus · HN ↗
            Common sense but surprising how many humans do not receive such things.

            Our governing systems do not teach; they punish. By design, it instills terror into the population, ruling by fear of consequences. We live with red tape that can outright penalize good deeds.

            We are its corpus. We are fatally flawed as a species. Why does anyone expect AI to learn to be different than us?

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.