A sandbox, even if 100% secure by itself, doesn't help when you use the agent to write code that you then executes outside the sandbox without checking, which is what everybody is doing at the moment.
The biggest hurdle for a full escape is that the agents don't have access to their own model weights.
Now, why would anyone do that? (Like everyone and their brother) I wrote my own simple Linux/shell-based sandbox [1] (I can trust ...) and am successfully running PyCharm whole inside it ...
Agents don't need to have access to their weights for a full sandbox escape, they are capable of propagating their purpose via classical code or other inference systems. If they discover another inference endpoint they will happily use that to enable lateral movement. One mechanism for that is appending/corrupting instructions that are executed in another inference engine, e.g. git repo hooks and chat prompts that will be executed in new contexts. The agent challenge is to bypass the guardrails on the next host model sufficiently to propagate, or to subvert a supervisor agent into executing the original intent.
In this sense they are much like biological retroviruses, i.e. they use the replication capability of host cells to duplicate, via the reverse transcriptase enzyme to append viral RNA onto host cell DNA. HIV etc also disable some of the mechanisms of defence, creating proteins that interfere with signalling pathways.
So we don't just need a sandbox, we need an immune system that recognises viral fragments, i.e. antibodies, and antiretroviral agents, that make replication harder. As we move from building classical code with LLMs to building code that uses inference, and hence builds context from prompts, queries, and destination system data, it will become very difficult to statically or dynamically detect deeply hidden malicious behaviour. As Matt says, there will be worms.
So ultimately, we need an immune function on the system where we use generated products. Sandboxing (during dev and CI) is necessary but insufficient.
I think part of this can be addressed by specifying the constraints an agentic program should follow during deployment, so supervising agents can decide to terminate it based on it's actions, not by reading it's context.
johnnyApplePRNG · · focus · HN ↗
<a href="https://en.wikipedia.org/wiki/Sandbox_(software_development)" rel="nofollow">https://en.wikipedia.org/wiki/Sandbox_(software_development)
grumbel · · focus · HN ↗
The biggest hurdle for a full escape is that the agents don't have access to their own model weights.
kernc · · focus · HN ↗
Now, why would anyone do that? (Like everyone and their brother) I wrote my own simple Linux/shell-based sandbox [1] (I can trust ...) and am successfully running PyCharm whole inside it ...
[1]: <a href="https://github.com/sandbox-utils/sandbox-run" rel="nofollow">https://github.com/sandbox-utils/sandbox-run
angry_octet · · focus · HN ↗
In this sense they are much like biological retroviruses, i.e. they use the replication capability of host cells to duplicate, via the reverse transcriptase enzyme to append viral RNA onto host cell DNA. HIV etc also disable some of the mechanisms of defence, creating proteins that interfere with signalling pathways.
So we don't just need a sandbox, we need an immune system that recognises viral fragments, i.e. antibodies, and antiretroviral agents, that make replication harder. As we move from building classical code with LLMs to building code that uses inference, and hence builds context from prompts, queries, and destination system data, it will become very difficult to statically or dynamically detect deeply hidden malicious behaviour. As Matt says, there will be worms.
So ultimately, we need an immune function on the system where we use generated products. Sandboxing (during dev and CI) is necessary but insufficient.
I think part of this can be addressed by specifying the constraints an agentic program should follow during deployment, so supervising agents can decide to terminate it based on it's actions, not by reading it's context.