‹ BackHN Continuity

Thread

Is sandboxing sufficient to contain rogue agents?

52 points · 99 comments · zdw

  1. Gigachad · · focus · HN ↗
    Seems to me that the problem is that if you sandbox agents enough to be safe, they can't do anything useful. And when you give them the tools to be useful, they can go off the rails in ways you didn't expect.

    Perhaps the answer is to have another agent who's goal is not to complete the given task, but to spot cheating or malicious behavior. We have seen some evidence that having AI review AI generated code actually does provide some value. You don't need a different model, just one which has been given the goal of finding flaws rather than achieving the task.

    1. mrweasel · · focus · HN ↗
      That does seem a little like solving the problems in AI by using more of it. I do see the idea, but if we're truly dealing with subversive agents on the level that the AI companies wants us to believe, then won't we need to deal with the first agent trying trick the second on?

      I still feel it would be much better to control the training data much more tightly. You'd still need agents with "hacking" abilities, for cyber security testing, but your average coding agent doesn't. So coding agents gets trained to be good citizens, respect autorisations, rejections, rate-limiting and so on.

      Sandboxing seems like a dead end for systems you inherently want to roam the internet and your file system.

      1. chrisjj · · focus · HN ↗
        So control training data to ensure good behaviour.

        I wonder how?

        Train on only stories of good deeds?

        On only works of good people?

        Or... what?

        1. mrweasel · · focus · HN ↗
          Mostly I was thinking good code. Exclude code that doesn't exits when encountering a 403, exclude code that doesn't have a back-off when encountering a 429.

          Teach the models that a 403 is you doing something you're not suppose to do, that is an existing status. There's only one action you're allowed to take on a 403 and that is to stop. No retry, no trying other API keys.

          The current approach with broad training and sandboxing to avoid misbehaviour isn't viable. It's much better to train the models to respect e.g. http status code and that they are not to be circumvented. Models for security research most obviously be trained differently.

          Smaller and more specialized models, with fewer, but targeted capabilities, seems to me to be a safer approach. If a model doesn't "know" that people leak API keys on Github, then it has no reason to go looking for them. If the current models are as "smart" as we're lead to believe, then guardrails and sandboxes aren't going to help, unless you lock the agents down to the point where they aren't useful. So dumb down the models.

          1. chrisjj · · focus · HN ↗
            > The current approach with broad training and sandboxing to avoid misbehaviour isn't viable.

            This.

            > It's much better to train the models to respect e.g. http status code and that they are not to be circumvented.

            I think you'd find respect requires intelligence, and is well out of scope of a next-token predictor.

            But I'm sure someone will try, and I will be interested to see.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.