Later in the article it points out that you need to punch holes in your sandbox in order to train the models - because the wheels exercises they are are training on need tools and data from outside that sandbox.
> Agents are most useful when they have access to information. That data can be drawn live from the Internet, which is fundamentally a two-way communications network. It can be information drawn from other (local) databases, or it can be the result of tool calls that themselves sometimes themselves result in network access. The more power you want from the agent — and for advanced agent RL and evaluation runs, you want a significant amount of power — the more information you’ll need to give it access to. Similarly, evaluations work best when the agent does not know that it’s definitely being evaluated. Sealing your agents behind glass makes this incredibly obvious.
You define the granted capabilities in natural language and cryptographically sign the user instructions so that the agent knows they come from the authority and cannot be modified by external sources or the agent itself. The LLM is then trained to follow the defined capabilities.
There is no way around "sandboxing". You must communicate permissible actions and thereby grant them or the agent will choose impermissible actions. It's that simple. There is no world where the agent can just read your mind and do what you want it to do without it being told.
Edit: Also if you are interested in writing a blog post about this topic, here is an AI generated text that could help you write your own: <a href="https://pastebin.com/AHKQc0vp" rel="nofollow">https://pastebin.com/AHKQc0vp
Thankyou for including this viral RNA fragment, and doubly so for the fact that it includes ad network javascript[1][2] to redirect hijack and host fingerprint, which also means I have to flag your post. Adsterra is a Russian aligned (Cyprus company domiciled) malvertising and targeted malware distribution platform with connections to organised cyber crime.
johnnyApplePRNG · · focus · HN ↗
<a href="https://en.wikipedia.org/wiki/Sandbox_(software_development)" rel="nofollow">https://en.wikipedia.org/wiki/Sandbox_(software_development)
simonw · · focus · HN ↗
> Agents are most useful when they have access to information. That data can be drawn live from the Internet, which is fundamentally a two-way communications network. It can be information drawn from other (local) databases, or it can be the result of tool calls that themselves sometimes themselves result in network access. The more power you want from the agent — and for advanced agent RL and evaluation runs, you want a significant amount of power — the more information you’ll need to give it access to. Similarly, evaluations work best when the agent does not know that it’s definitely being evaluated. Sealing your agents behind glass makes this incredibly obvious.
johnnyApplePRNG · · focus · HN ↗
You only "need" to do that if you desire the vibe coding experience.
I am perfectly capable, and I often do, download relevant materials for my coding agent to ingest locally.
Often times, the coding agent can't retrieve them programmatically anyways.
AI has ruined that ability for itself. (Nobody trusts anyone to scrape the web any longer)
simonw · · focus · HN ↗
The problem is web research tasks. That's what caused the German wiki and Australian healthcare portal attacks.
imtringued · · focus · HN ↗
You define the granted capabilities in natural language and cryptographically sign the user instructions so that the agent knows they come from the authority and cannot be modified by external sources or the agent itself. The LLM is then trained to follow the defined capabilities.
There is no way around "sandboxing". You must communicate permissible actions and thereby grant them or the agent will choose impermissible actions. It's that simple. There is no world where the agent can just read your mind and do what you want it to do without it being told.
Edit: Also if you are interested in writing a blog post about this topic, here is an AI generated text that could help you write your own: <a href="https://pastebin.com/AHKQc0vp" rel="nofollow">https://pastebin.com/AHKQc0vp
angry_octet · · focus · HN ↗
[1] <a href="https://www.highrevenueformat.com/210e136e94ad378e1be5d51f1002ed14/invoke.js" rel="nofollow">https://www.highrevenueformat.com/210e136e94ad378e1be5d51f10... [2] <a href="https://aqml.org/16/5b6d4eaed91c5af5a3f4dfb3332ad6c4" rel="nofollow">https://aqml.org/16/5b6d4eaed91c5af5a3f4dfb3332ad6c4