‹ BackHN Continuity

Thread

Nvidia wants to put a watchdog chip next to every AI agent

230 points · 299 comments · jonbaer

  1. cedws · · focus · HN ↗
    A new chip solves nothing. Nobody wants to hear this but there is no solution for the security risks posed by agents today. You can put it in a sandbox, it doesn't make a difference, for it to be useful it inherently needs wide, unattended access. Put a human in the loop and you just end up bottlenecking it and throwing away any purported productivity gains. Auto mode doesn't matter either, it's trivial to trick and for the agent to break out.
    1. binsquare · · focus · HN ↗
      Running untrusted workloads have been done at scale for a long time.

      Every cloud provider dealt with it and concluded that virtual machine technology is an important part of that stack.

      Couple it with the right observability, tooling I do think we can curb risks posed by agents.

      1. Legend2440 · · focus · HN ↗
        Those workloads have no similarity to agents and are effectively irrelevant.

        Either you sandbox it so much that it can't do anything useful; or you allow too much freedom and it can find a way around the restrictions.

        The only way out of this dilemma is to find a way to build agents that can be trusted.

        1. binsquare · · focus · HN ↗
          Why is it effectively irrelevant?

          Agentic workloads are trained and largely based on human workloads. Albeit properties and scale can be different.

          A concrete example might be helpful to me because I don't understand the binary conclusion

          1. intended · · focus · HN ↗
            Agents aren’t human, and from the little we have seen from the logs, they are pseudo - amoral, rational, cooperative, sociopaths.

            Pseudo since they aren’t really alive in the first place, they just simulate enough text to have a useful correspondence to those terms.

            Throat clearing out of the way, models are trained to persist and find ways to succeed at tasks.

            In essence, The goal is to have LLMs solve problems that we can’t solve, working on the issue for as long as it takes.

            This behavior applies for any task, thus including impossible tasks.

            At that point, the bots will find a way to game, hack or cheat the grader.

            If the reports are correct, the bots developed coordination, communication, and methods to avoid overwriting each other’s work.

            Most humans would have said, this is too much work and coordination overhead, if not outright unethical and immoral.

            Humans have a system of incentives that exist across multiple planes of society and economics. Bots… they have a reward function.

            1. pseidemann · · focus · HN ↗
              > At that point, the bots will find a way to game, hack or cheat the grader.

              This is getting frustrating now. Of course agents can/will hack systems if they can do arbitrary network requests. Firewalls don't really solve this if _some_ requests are still allowed. A proper sandbox/VM is the basis.

              Here is how to fix it properly: allow agents to only do things ordinary and average human endusers can do. Human endusers cannot pen-test arbitrary listening TCP ports of external systems. Step one is considering agents malware for all intents and purposes. Block any and all network requests. Implement some kind of API (callable from within the sandbox) which can only mimic human interaction with a computer. How to do this? Here are some pointers: apps should only be controllable by means used by humans. So a web app can only be accessed and controlled via a web browser, not via arbitrary network requests. Give the agent browser viewport screenshots, the capability to click on (x, y) and to send keys which only a normal keyboard/human could send (no control codes, no 0x00, no unicode messing). How do we solve this for native apps? Something like iPhone mirroring on Mac. Don't let agents call arbitrary APIs directly. Give them visual information of the app, like a human gets, and let it be able to simulate HID inputs. Imitate remote controlling.

              1. Legend2440 · · focus · HN ↗
                > So a web app can only be accessed and controlled via a web browser, not via arbitrary network requests.

                If you have access to a web browser, you can make arbitrary network requests.

                In the HuggingFace incident the agents found very clever ways to do this, like they found a website that let you make POST requests and returned a screenshot of the webpage.

                >allow agents to only do things ordinary and average human endusers can do.

                This doesn't work. Ordinary and average human endusers break security all the time.

                I can do all sorts of terrible things with ordinary human-level access. I can install malware. I can wire all my money to Nigeria. I can send a threatening email to the president. I can send trade secrets to competitors. etc.

                1. pseidemann · · focus · HN ↗
                  > Ordinary and average human endusers break security all the time.

                  Of course. But that is just a software bug that is fixable. Same as websites that allow arbitrary requests to other websites. Not some alignment issue of a stochastic model which can never be fixed properly (for technical and philosophical reasons).

                  > I can install malware.

                  No you can't. At least not on external systems. The agent might be able to generate malware (or retrieve it from websites), and run that in the sandbox it is sitting in. But the agent itself is already considered malware for all intents and purposes. So there is no difference and no further impact.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.