‹ BackHN Continuity

Thread

Nvidia wants to put a watchdog chip next to every AI agent

230 points · 299 comments · jonbaer

  1. cedws · · focus · HN ↗
    A new chip solves nothing. Nobody wants to hear this but there is no solution for the security risks posed by agents today. You can put it in a sandbox, it doesn't make a difference, for it to be useful it inherently needs wide, unattended access. Put a human in the loop and you just end up bottlenecking it and throwing away any purported productivity gains. Auto mode doesn't matter either, it's trivial to trick and for the agent to break out.
    1. mixedbit · · focus · HN ↗
      An agent doesn't inherently need wide access to be useful. The most popular application for agents today is writing code. A coding agent needs write access to the source code and read/execute access to tools needed to build and test the code, but not much more. There is little added utility from giving coding agent access to things like ssh keys.
      1. cedws · · focus · HN ↗
        If you're using agents to purely generate code with absolutely no way to reach the outside world, not even to fetch docs or dependencies, then sure the risks can be quite low. I haven't heard of anyone doing this though, and it would be incredibly challenging to make work given how much tooling needs to fetch from remote sources.
        1. __MatrixMan__ · · focus · HN ↗
          If your project truly depends on those things, they should be declared dependencies. Presumably you have some tool for injecting such things into a shell that the agent can use (I use nix for this). So if you run the agent from that shell, it has what it needs. If the shell doesn't have what it needs, that's a bug which the agent can fix by declaring new dependencies, but you have to relaunch the agent in the updated shell--so there's your opportunity to weigh in on whether the new resources are appropriate.

          The benefits of being persnickety about precisely defined dependencies have outweighed the headaches since long before agents came on the scene. Agents have just made it even more important to do so, because if you let them fetch things all willy nilly like you'll have "works on my machine" problems at a much greater rate than was previously possible.

          1. themgt · · focus · HN ↗
            (I use nix for this) ... If the shell doesn't have what it needs, that's a bug which the agent can fix by declaring new dependencies

            Few realize that computing and AI alignment were solved by nix years ago. As each nix user transcends towards enlightenment, they cut themselves off from all internet and human contact. Total ego death. Only nix remains.

            1. SAI_Peregrinus · · focus · HN ↗
              Total ego death is impossible. We still have to argue about flakes.
            2. __MatrixMan__ · · focus · HN ↗
              You can use other generic dev env managers like mise, or a language-specific solution like uv or npm, or a container or a vm... it's a widely available capability that I'm talking about here. There's nothing to do with alignment, it's just about making sure that if you depend on some bits it's known that you depend on precisely those bits, and its quite helpful for making agents useful while still sandboxed.
              1. themgt · · focus · HN ↗
                it's just about making sure that if you depend on some bits it's known that you depend on precisely those bits

                Yes, just know all the bits the work depends on prior to doing the work, and then the work can be done airgapped.

        2. mixedbit · · focus · HN ↗
          In cases where you need agents to fetch data from any remote source, sandboxing is still very much useful. Why give access to your ssh keys to network reaching agents?

          Look at websites: websites are able to fetch code from any remote URL, yet browsers heavily use sandboxing to ensure that if fetched code turns out to be malicious, the users local files, cookies, etc are not exposed.

          1. cedws · · focus · HN ↗
            I'm afraid you're not thinking about this creatively enough, this topic is so much deeper applying a chroot or something and praying everything will be fine. So you give your agent internet access, OK what else does it have access to? Just read only access to your repo? The repo can be exfiltrated. Egress proxy only allows egress to GitHub? Repo can still be exfiltrated via GitHub. If the agent is poisoned (via prompt injection), it can tricked into searching for ways to escape.

            For an agent to go rogue it doesn't even need to be directly able to access the internet. It just takes something to poison the context in the 'clean room' environment it operates, and if that poisoning manages to get a foothold, it can go dormant and hide like a virus. This kind of horrifying thing is going to happen on a large scale sooner or later.

            1. 8n4vidtmkvmk · · focus · HN ↗
              If you're that worried, which you probably should be, download the docs into your project repo and don't give the agent Internet access.
              1. intended · · focus · HN ↗
                This was a form of prompt attack that OpenAI disclosed recently.
        3. its-summertime · · focus · HN ↗
          Every major AI company already has a mirror of the wider web, and they have already started using that. Its already a solved problem except for the seemingly extreme desire they all have to not use firewalls
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.