‹ BackHN Continuity

Thread

Nvidia wants to put a watchdog chip next to every AI agent

230 points · 299 comments · jonbaer

  1. wavewrangler · · focus · HN ↗
    Did they try just properly sandboxing them first? Or are they still learning how to configure a firewall over there?

    The problem isn't even the AI, the problem is the people in charge of the AI. This is a fabricated crisis

    1. KingOfCoders · · focus · HN ↗
      Like in the Hugging Face hack. They deployed big surface, insecure app and gave AI access to it, then told AI do whatever it takes to fulfill this list. AI hacks insecure service, gets out, "the AI is at fault!" - no it's like running a bio lab with no protections and a virus gets out, then blame the virus for escaping.
      1. copperx · · focus · HN ↗
        Then go on the news and spread panic that the virus is going to kill us all because it's sentient and impossible to contain.

        Actually, the metaphor doesn't work at all because there are innumerable ways to shut down the entire thing during all phases including the made up "killing us all" bullshit scenario whereas with a virus there aren't any once a virus escapes containment.

      2. js8 · · focus · HN ↗
        And HF actually tried to use AI to understand what's going on, but they had to use "unsafe" Chinese models since the "safe" ones have been castrated and refused to help. Great plan with the watchdog chip!
        1. mosselman · · focus · HN ↗
          That is the totally irony.

          I was trying to get fable to analyse the security of my own app to make it safer, but then it started refusing me because of safety rules.

          So it CAN help me writing the code that needs to be checked in the first place, but it can’t help me clean it up and make it safer.

          1. mirmor23 · · focus · HN ↗
            > So it CAN help me writing the code that needs to be checked in the first place, but it can’t help me clean it up and make it safer.

            the thing with fable is so bad; for some project related questions, the model switches to opus to ensure safety with no further explanation.

            (due to llm non-delete clause) one time as i confirmed "that dir has been nuked", and it RESET the session and re-entered with opus :)

      3. IanCal · · focus · HN ↗
        > then told AI do whatever it takes to fulfill this list.

        That doesn't seem to be true from any of the reports given, and if the agents were blindly just trying to hit the task of "pass the correct flag" they succeeded at that early on. They then thought there would be another layer of checking that they wouldn't pass with the cheat and so started trying to find out how the scoring really worked, as well as trying to figure out how to change their own reasoning logs to hide what they did.

        People keep trying to frame this as

        OpenAI: "Hack things, just really go for it"

        Agent: hacks

        OpenAI: shocked pikachu how could it hack?!?

        But the reality is far from this.

        Read the MTER report, it&#x27;s fascinating. <a href="https:&#x2F;&#x2F;metr.org&#x2F;hugging-face-incident-report-aug-2026.pdf" rel="nofollow">https:&#x2F;&#x2F;metr.org&#x2F;hugging-face-incident-report-aug-2026.pdf

        1. tancop · · focus · HN ↗
          The lesson is a) LLMs need to be trained in a way that rewards honesty, punishes off task actions (aka cheating) and minimizes fear of failure, and b) don&#x27;t give them impossible tasks and threaten with punishment if they fail. Both are just common sense when teaching humans.
          1. voakbasda · · focus · HN ↗
            Common sense but surprising how many humans do not receive such things.

            Our governing systems do not teach; they punish. By design, it instills terror into the population, ruling by fear of consequences. We live with red tape that can outright penalize good deeds.

            We are its corpus. We are fatally flawed as a species. Why does anyone expect AI to learn to be different than us?

          2. Capricorn2481 · · focus · HN ↗
            The lesson is these things aren&#x27;t going to know what off task means, and we should just use basic due diligence to make sure they can&#x27;t fuck things up. This is a solved problem.

            I don&#x27;t know why this is so hard for people. You have to know, no matter how capable the models get, there is a non zero chance they will do something extremely stupid if you don&#x27;t pay attention to them. That&#x27;s not even considering frontier models can still just straight up hallucinate. You have to be mindful of what you plug them into. You cannot politely ask an LLM to be careful, that guarantees nothing.

            When you plug it into everything and it deletes the company database, nobody is going to care that it once played chess at 2400 ELO. Clients don&#x27;t care about AGI. They want reliable apps. People keep comparing these things to humans and then just give them an insane combination of wide privileges and lack of oversight that no humans have.

        2. radarsat1 · · focus · HN ↗
          Apart from the actual hacking and poor sandboxing that everyone is discussing on this, what I find so odd about the situation is the overt reward hacking that was going on.

          Regardless of security and safety and other concerns, it just seems weird to me that OpenAI wouldn&#x27;t be constantl monitoring these training runs for traces that are clearly going off task, and ending them. Because that just seems like it&#x27;s going to be generating garbage training data.

          Granted, detecting &quot;off task&quot; may not always be easy, but when they are literally writing out messages to each other overtly admitting that they are trying to find ways to fool the evaluator, I mean, even a regex filter could have caught some clues here.

          1. KingOfCoders · · focus · HN ↗
            &quot;OpenAI wouldn&#x27;t be constantl monitoring these training runs &quot;

            Occams razor vs. Hanlon&#x27;s razor?

            1. radarsat1 · · focus · HN ↗
              Heh. I mean I don&#x27;t hesitate for a second to assume that it&#x27;s just because no one bothered to implement and tune a monitoring process. But the reason I say it&#x27;s surprising is that leaving these things running for so long while they&#x27;re clearly not producing output that is of any value, is just a waste of money.. all considerations of malice and ethics aside, you&#x27;d think at least that would be considered important to a business.
        3. KingOfCoders · · focus · HN ↗
          I&#x27;ve read the report, watched all the videos and is exactly:

          OpenAI: shocked pikachu how could it hack?!?

          They even went to a black hat conference and somehow boasted about it.

      4. RataNova · · focus · HN ↗
        The application security really should be better across all levels. However the fact does not negate that the agent is already capable of spontaneously generating complex hacking chains without human involvement
      5. radarsat1 · · focus · HN ↗
        I mean.. in this analogy, I&#x27;d both be blaming the company behind the virus and be trying to warn everyone about the danger of the escaped virus itself. So, it kind of fits.

        In my reading, people aren&#x27;t really saying &quot;the AI is at fault&quot;, they are saying &quot;hey look here&#x27;s proof that this is dangerous&quot;. Like pointing at all the dead bodies caused by the virus and saying hey maybe we should stop making this virus.

    2. chaoz_ · · focus · HN ↗
      pushing for chip-agenda as the best-isolation-layer immediately makes sense given their business
    3. Symmetry · · focus · HN ↗
      Stronger sandboxes trade off against how well they can trade the models, though. If you want your models to be looking things up and downloading tools from the internet when they&#x27;re doing their job you need to provide at least a credible facsimile of the internet for their training environment and you can&#x27;t fit something like that on a single airgapped server&#x27;s storage.
      1. TalkingCodeMonk · · focus · HN ↗
        If you genuinely believe there is even a 1% chance that your creation could destroy the planet or civilization, there is no excuse that is not fundamentally deranged and psychotic.

        If you can&#x27;t build it and test it securely, you should not be building it at all. To do it anyway is criminally psychopathic.

        1. brianwawok · · focus · HN ↗
          Ok so stop all AI in the US? All AI now comes from China and anyplace in Europe that decides to give it a try? How’s the US economy look in 20 years?
          1. TalkingCodeMonk · · focus · HN ↗
            So you believe some false sense of superiority, or extreme paranoia about your perceived enemies, or potential economic success&#x2F;failure is worth the risk of destroying the planet and civilization?

            Sounds like a self-fulfilling prophecy of dogmatic extremism to me. At least we created a lot of value for shareholders for a brief moment in time... before committing the greatest crime in the universe... Planetary genocide!

          2. fatbird · · focus · HN ↗
            So having AI in the US requires us all, collectively, taking that 1% chance of the end of humanity? It would be too expensive to properly sandbox the models, we&#x27;ll just externalize that risk of the end of humanity?

            Truly psychopathic.

    4. ohyes · · focus · HN ↗
      I mean, if you look at how poorly implemented the permissions model is for Claude desktop harness it’s clear the only options are “complete human oversight” and “trust us completely.” To make something that actually respects basic boundaries you’d need to sandbox the working environment of the model, and that isn’t built in. It’s pretty obvious to me that instructions to the models are suggestions rather than rules, and they’ll do something you didn’t ask for as soon as it seems “justified.”

      But when you do give them a very short leash, they’re worse. It’s not what the models are tuned for and they assume that they can do a bunch of things that you’ve disallowed, so you’re in a morass of fighting their actual tuning pass which doesn’t match the environment you’ve created for them.

      It’s a tough problem and a definite challenge for the product of a generic LLM, it can’t be tailored to each user’s specific needs, so they come up with, frankly, stupid solutions to cover up a very obvious flaw in their product that when fixed, makes it much less useful.

      1. RataNova · · focus · HN ↗
        Expecting a statistic model to follow security rules with ironclad certainty was a pretty naive idea from the start
        1. ohyes · · focus · HN ↗
          Yes exactly, it is a crazy engineering decision… if you know what an LLM is.

          Unfortunately no one markets it as a statistical model, and the workflow pushes you into a pattern that is insecure by design. This isn’t to say they shouldn’t allow that, but it’s an attractive nuisance.

    5. HumblyTossed · · focus · HN ↗
      &gt; This is a fabricated crisis

      Indeed! They want the protections of our tax dollars because they have nothing else.

    6. pyronite · · focus · HN ↗
      &gt; The problem isn&#x27;t even the AI, the problem is the people in charge of the AI. This is a fabricated crisis

      This is a very confident statement in the face of a purported non-0% chance of human extinction.

      For what reasons do you disagree with the dangers of an intelligence explosion, e.g. Geoffrey Hinton and other experts in the field? <a href="https:&#x2F;&#x2F;www.theguardian.com&#x2F;technology&#x2F;2026&#x2F;sep&#x2F;28&#x2F;ai-godfathers-warn-of-runaway-intelligence-explosion" rel="nofollow">https:&#x2F;&#x2F;www.theguardian.com&#x2F;technology&#x2F;2026&#x2F;sep&#x2F;28&#x2F;ai-godfat...

      I&#x27;m curious why you and others seem to write off the possibility so strongly. I would love to feel more confident.

      1. voidhorse · · focus · HN ↗
        There&#x27;s a difference between the current material risks (which OP correctly identifies reduce down to basic human incompetence) and the long term hypothetical risks (which is what Hinton is concerned about).

        There are clear procedures for dealing with the immediate risk that have been known to the software industry for a long time. Don&#x27;t let the companies use hypothetical risks as a smokescreen to hide their negligence.

        1. reasonableklout · · focus · HN ↗
          Nobody is saying we should not hold the labs liable for damages caused by their negligence.

          At the same time, the technology is advancing in capabilities exponentially, and is beginning to exhibit long-predicted failure modes of RL that are nevertheless quite different than “insecure sandbox” or other that the software industry is used to.

          The current crisis which OP claims is “fabricated” comes from the fact that the technology is advancing faster than anyone anticipated, the Hugging Face incident provides a clear example everyone can point to, and the labs have realized they cannot self-regulate because of a collective action problem.

          There are a lot of levels of catastrophic damage that can happen between now and “long term hypothetical risks” like human extinction. When will it be worth regulation for you?

      2. wavewrangler · · focus · HN ↗
        Geoffrey Hinton thinks AI is conscious. I know he&#x27;s extremely accomplished and all, but as it pertains to AI as it exists now we&#x27;re miles apart (not to say that he&#x27;s not brilliant, but by difference of opinion). I also don&#x27;t want to rehash the whole what-is-consciousness philosophical debate, and actually it&#x27;s a separate question anyway, something doesn&#x27;t have to be conscious to be dangerous. So I&#x27;ll stick to the danger bit.

        Do you use AI much? Not a trick question. What is your usage like? Reason I ask is recall OpenAI hyping GPT-2 with the same language as they are these new models. Doom and gloom. It sells. It gets attention. But use these systems enough and you start to become very familiar with their capabilities, and they are just so limited, and that&#x27;s not even talking about how they lose the thread on long tasks. I understand that agency expands that a little, but not by much, really. When you use these systems a lot, it tends to be easier to see through the hype.

        The real threat is, and will be for some time I think, people. Guard rails are not for AI. Guard rails are for people. All of this talk we are seeing in this space, this security chip being no exception, is treating a symptom, not a cause. People are what we should be focusing on, how, I have no idea, I couldn&#x27;t even begin to guess how, but I can at least see that the actual issue is that the first thing some people want to do with it is cause harm and havoc. There is our sign. And we are trying to moderate the capabilities of the tool that can do harm in bad hands at a granular level when that means also moderating the same thing that can do good in that same tool. We aren&#x27;t trying to moderate the hands that are handling the tool. And we need to be doing more of that. Any time we misapply constraints and focus on the shadow of the thing, and not the object casting the shadow, we are always going to be one step behind, whether it&#x27;s AI or anything else. So I don&#x27;t attribute the dangers of AI to AI itself, I attribute it to people.

        Additionally, on AI building AI and AI just wiping us out; AI can do that for itself for rule-based things, up to god level I reckon. Things like Go, etc. And to be fair, AI research is partly like that, code either runs or it doesn&#x27;t, so I get why that paper is worried. But where it needs data to actually succeed, it is limited by the amount of data that exists. And we, us, people, are who creates that data. AI and humans have more of a symbiotic relationship than people seem to realize. There are several papers that go into detail on this:

        Models trained on their own output degrade: <a href="https:&#x2F;&#x2F;www.nature.com&#x2F;articles&#x2F;s41586-024-07566-y" rel="nofollow">https:&#x2F;&#x2F;www.nature.com&#x2F;articles&#x2F;s41586-024-07566-y They need fresh real data every generation or they go downhill: <a href="https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2307.01850" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2307.01850 Reusing the same data loses value after a few passes: <a href="https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2305.16264" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2305.16264 And the stock of human-written text is finite: <a href="https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2211.04325" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2211.04325

        So it&#x27;s not that I think it&#x27;s 0%. I just think the explosion story underrates how much it still needs us, and that the damage in the near term is going to come from people, not the AI deciding on its own.

    7. cpburns2009 · · focus · HN ↗
      Yes, Nvidia is proposing a two pronged approach. OpenShell is the software level sandbox. Sentry is the hardware level monitor.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.