‹ BackHN Continuity

Thread

Is sandboxing sufficient to contain rogue agents?

52 points · 99 comments · zdw

  1. Gigachad · · focus · HN ↗
    Seems to me that the problem is that if you sandbox agents enough to be safe, they can't do anything useful. And when you give them the tools to be useful, they can go off the rails in ways you didn't expect.

    Perhaps the answer is to have another agent who's goal is not to complete the given task, but to spot cheating or malicious behavior. We have seen some evidence that having AI review AI generated code actually does provide some value. You don't need a different model, just one which has been given the goal of finding flaws rather than achieving the task.

    1. bigstrat2003 · · focus · HN ↗
      If you can't trust a tool, you shouldn't be running it at all. It's really quite simple. It doesn't matter how useful it is if you can't actually have confidence in using it safely.
      1. Gigachad · · focus · HN ↗
        People will use the tool regardless. so it’s a race to try to make it safe before something truely bad happens.
      2. dipper139 · · focus · HN ↗
        I don't think it's about trust but rather incomplete evaluation. Evaluating the model on its capacity to refuse a task or to question its prompt is something recent when you look at it, i feel current AI is really just an immature solution and we are just yet realizing the mistakes that have been made for so long
      3. rlpb · · focus · HN ↗
        And yet we we all use human written software even though we can be confident that the next severe software vulnerability to be found in it is just round the corner.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.