‹ BackHN Continuity

Thread

Nvidia wants to put a watchdog chip next to every AI agent

230 points · 299 comments · jonbaer

  1. wavewrangler · · focus · HN ↗
    Did they try just properly sandboxing them first? Or are they still learning how to configure a firewall over there?

    The problem isn't even the AI, the problem is the people in charge of the AI. This is a fabricated crisis

    1. KingOfCoders · · focus · HN ↗
      Like in the Hugging Face hack. They deployed big surface, insecure app and gave AI access to it, then told AI do whatever it takes to fulfill this list. AI hacks insecure service, gets out, "the AI is at fault!" - no it's like running a bio lab with no protections and a virus gets out, then blame the virus for escaping.
      1. IanCal · · focus · HN ↗
        > then told AI do whatever it takes to fulfill this list.

        That doesn't seem to be true from any of the reports given, and if the agents were blindly just trying to hit the task of "pass the correct flag" they succeeded at that early on. They then thought there would be another layer of checking that they wouldn't pass with the cheat and so started trying to find out how the scoring really worked, as well as trying to figure out how to change their own reasoning logs to hide what they did.

        People keep trying to frame this as

        OpenAI: "Hack things, just really go for it"

        Agent: hacks

        OpenAI: shocked pikachu how could it hack?!?

        But the reality is far from this.

        Read the MTER report, it&#x27;s fascinating. <a href="https:&#x2F;&#x2F;metr.org&#x2F;hugging-face-incident-report-aug-2026.pdf" rel="nofollow">https:&#x2F;&#x2F;metr.org&#x2F;hugging-face-incident-report-aug-2026.pdf

        1. tancop · · focus · HN ↗
          The lesson is a) LLMs need to be trained in a way that rewards honesty, punishes off task actions (aka cheating) and minimizes fear of failure, and b) don&#x27;t give them impossible tasks and threaten with punishment if they fail. Both are just common sense when teaching humans.
          1. voakbasda · · focus · HN ↗
            Common sense but surprising how many humans do not receive such things.

            Our governing systems do not teach; they punish. By design, it instills terror into the population, ruling by fear of consequences. We live with red tape that can outright penalize good deeds.

            We are its corpus. We are fatally flawed as a species. Why does anyone expect AI to learn to be different than us?

          2. Capricorn2481 · · focus · HN ↗
            The lesson is these things aren&#x27;t going to know what off task means, and we should just use basic due diligence to make sure they can&#x27;t fuck things up. This is a solved problem.

            I don&#x27;t know why this is so hard for people. You have to know, no matter how capable the models get, there is a non zero chance they will do something extremely stupid if you don&#x27;t pay attention to them. That&#x27;s not even considering frontier models can still just straight up hallucinate. You have to be mindful of what you plug them into. You cannot politely ask an LLM to be careful, that guarantees nothing.

            When you plug it into everything and it deletes the company database, nobody is going to care that it once played chess at 2400 ELO. Clients don&#x27;t care about AGI. They want reliable apps. People keep comparing these things to humans and then just give them an insane combination of wide privileges and lack of oversight that no humans have.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.