‹ BackHN Continuity

Thread

Is sandboxing sufficient to contain rogue agents?

52 points · 99 comments · zdw

  1. Luker88 · · focus · HN ↗
    I tried using opencode permissions to limit agents.

    It it completely pointless. you can't even make a "read-only" agent. allow "cat *" for every file? congratulation, that allows "cat file > output" and now you have read write.

    Allow python? more free reign that allowing all bash. The models (qwen or claude) will still try to use the disallowed things multiple times.

    read/edit permission are bad enough that the model themselves don't understand why they don't have permissions: they double check the conf, and think they should have access.

    I am switching to using one firejail per project to containerize as much as possible, and leave all permissions to allow.

    I have no idea how to limit network access, and I have no idea how to prompt and steer subagents when they are going off the rails.

    The whole thing is built to be completely impossible to limit and steer.

    1. chrisjj · · focus · HN ↗
      In a simple local data processing task, I told ChatGPT to stop searching web. It agreed, then continued regardless - referring me to an OpenAI help page saying it may need to search the web.

      > The whole thing is built to be completely impossible to limit and steer.

      And by people completely impossible to limit and steer.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.