A trivial bypass I can imagine is spawning a new process via `ssh localhost foo` - the new process forks from sshd, not the client. This is simple enough that an LLM could come up with it all on its own, if it feels that you're getting in its way.
It's a broad class of bypasses that can apply to just about any "long running daemon can be instructed to spawn a new child" situation.
Yeah, you're right. This tool was something I wrote to control unintentional leaks that I observed when working on something at my internship. While writing this I did try accounting for things like ssh, systemd-run and using D-Bus to run other processes but I wasnt really able to think of a way to stop this type of leak.
Also there's not much I can do to stop a dedicated LLM from getting out if its intention is to escape containment beyond maybe like a sandbox with no network access and access to these daemons. I appreciate the comment though, I really should update the readme to mention this.
The leaks likely happened in the first place because the model forgot it had a tool call for running properly managed background tasks (claude seems to do this all the time). I'm not saying a the model is going to be malicious and try to escape, I'm just saying it's going to employ workarounds to get its job done.
Retr0id · · focus · HN ↗
It's a broad class of bypasses that can apply to just about any "long running daemon can be instructed to spawn a new child" situation.
CG144 · · focus · HN ↗
Also there's not much I can do to stop a dedicated LLM from getting out if its intention is to escape containment beyond maybe like a sandbox with no network access and access to these daemons. I appreciate the comment though, I really should update the readme to mention this.
Retr0id · · focus · HN ↗