‹ BackHN Continuity

Thread

Is sandboxing sufficient to contain rogue agents?

52 points · 99 comments · zdw

  1. Gigachad · · focus · HN ↗
    Seems to me that the problem is that if you sandbox agents enough to be safe, they can't do anything useful. And when you give them the tools to be useful, they can go off the rails in ways you didn't expect.

    Perhaps the answer is to have another agent who's goal is not to complete the given task, but to spot cheating or malicious behavior. We have seen some evidence that having AI review AI generated code actually does provide some value. You don't need a different model, just one which has been given the goal of finding flaws rather than achieving the task.

    1. baxtr · · focus · HN ↗
      That could work.

      My thinking is: If AI is really smart, AGI smart for some, why wouldn't it be able to understand - over time - what is appropriate and what not?

      Maybe we need more human intervention to train it properly. Maybe we need constant intervention by a "police" agent.

      1. ben_w · · focus · HN ↗
        A problem is the agents who hacked Hugging Face already understood (we can tell because they wrote it down) that their actions were not appropriate, and then did those things anyway.

        "Helpful, harmless, honest": we can even ignore "honest" for this point, for tasks like the HuggingFace incident (ExploitGym with impossible challenges), we can pick anywhere on the spectrum from "helpful" to "harmless", the former being "completing the task" the latter being "refusing because completion required unlawful behaviour".

        (The agents in that case were also not "honest" in this case; this is an extra problem, and does not invalidate how helpful-vs-harmless is already a tradeoff).

        1. dns_snek · · focus · HN ↗
          > already understood (we can tell because they wrote it down)

          No, generating tokens doesn't equal understanding. Does GPT-2 understand human emotions just because it can generate some text talking about them?

          1. cindyllm · · focus · HN ↗

            [dead]

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.