Is sandboxing sufficient to contain rogue agents?
Thread
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
Is sandboxing sufficient to contain rogue agents?
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
varman11 · · focus · HN ↗
[dead]
Gigachad · · focus · HN ↗
Perhaps the answer is to have another agent who's goal is not to complete the given task, but to spot cheating or malicious behavior. We have seen some evidence that having AI review AI generated code actually does provide some value. You don't need a different model, just one which has been given the goal of finding flaws rather than achieving the task.
bigstrat2003 · · focus · HN ↗
Gigachad · · focus · HN ↗
dipper139 · · focus · HN ↗
rlpb · · focus · HN ↗
mrweasel · · focus · HN ↗
I still feel it would be much better to control the training data much more tightly. You'd still need agents with "hacking" abilities, for cyber security testing, but your average coding agent doesn't. So coding agents gets trained to be good citizens, respect autorisations, rejections, rate-limiting and so on.
Sandboxing seems like a dead end for systems you inherently want to roam the internet and your file system.
msdz · · focus · HN ↗
> That does seem a little like solving the problems in AI by using more of it
Yes, and IIRC Google used this as part of a technique against prompt injection already [0], back when models were way more susceptible to it.
[0] Cf. CaMeL: <a href="https://arxiv.org/abs/2503.18813" rel="nofollow">https://arxiv.org/abs/2503.18813
chrisjj · · focus · HN ↗
mrweasel · · focus · HN ↗
chrisjj · · focus · HN ↗
This.
> It's much better to train the models to respect e.g. http status code and that they are not to be circumvented.
I think you'd find respect requires intelligence, and is well out of scope of a next-token predictor.
But I'm sure someone will try, and I will be interested to see.
aytigra · · focus · HN ↗
LoganDark · · focus · HN ↗
Such a model doesn't yet exist, of course.
saagarjha · · focus · HN ↗
LoganDark · · focus · HN ↗
cassianoleal · · focus · HN ↗
hanibrel · · focus · HN ↗
[dead]
baxtr · · focus · HN ↗
My thinking is: If AI is really smart, AGI smart for some, why wouldn't it be able to understand - over time - what is appropriate and what not?
Maybe we need more human intervention to train it properly. Maybe we need constant intervention by a "police" agent.
mulmen · · focus · HN ↗
mdp2021 · · focus · HN ↗
Ann's intelligence and Bob's morality will seem orthogonal. Charles' morality is a function of C.'s intelligence as an ability as an effort spent to reach the current moral conclusion.
mulmen · · focus · HN ↗
mdp2021 · · focus · HN ↗
mulmen · · focus · HN ↗
dns_snek · · focus · HN ↗
mdp2021 · · focus · HN ↗
Where is the argument? If Bob has determined that «preserving life on earth» has some important weight, for Bob's there unspecified own reasons, and has also determined that the best course of action would be «to eradicate the human species», the one question is whether Bob is right or not. What was stated is, that Ethical Calculus is a function of Intelligence - of course it is, it is a structure of assessments.
> A century ago some Bobs decided
And who has told you that those "bobs" were "intelligent"?!?!?!
> So no
All you have proven is that you dislike some moral conclusion of some decisors. Which is trivial, obvious, and part of the already stated framework - proper ethical judgement requires proper general judgement (Intelligence).
mulmen · · focus · HN ↗
mdp2021 · · focus · HN ↗
I have never said that. That is just your reconstruction.
There is Decision Theory. It outputs optimal action through evaluation of strategies, of contexts, of principles. In order to properly get the strategies, the contexts, the principles, you need that skill that approximates ideas to truth - and such skill is named Intelligence. Optimal action decided in light of principles within a well assessed context is ethical behaviour. Ethical behaviour hence requires Intelligence.
As written, «Ethical Calculus is a function of Intelligence - of course it is, it is a structure of assessments».
> possible to act morally without being intelligent
Random correct behaviour proves nothing. Of course one can guess the roll of a dice roughly every sixth event. If you behave "well" but do not know why that is "well", that is like guessing. "Good" behaviour without intellectual awareness is like memorizing arithmetic (multiplication tables) without knowing why those memorized notions are correct.
> possible to be intelligent without acting morally
No, because by definition that would be a fault in Intelligence. If your action was imperfect, suboptimal, it is because you could not think of a better action or understand that the other action was better. If two choices C1 and C2 can be ranked, there is a reason for their order; knowing and understanding that reason is the task of Intelligence.
mulmen · · focus · HN ↗
mulmen · · focus · HN ↗
mdp2021 · · focus · HN ↗
Why does not the NN or whatever entity «understand ... what is appropriate»? Because it is not intelligent enough! How does an entity know what is moral and what is an atrocity? By being intelligent enough! Could an entity act appropriately without knowing it? Yes, it happens all the time if the wind blows right, but we do not rely on that! How to make something act appropriately? Well, an Intelligent entity in the loop must be there to know what is appropriate and what not!
dns_snek · · focus · HN ↗
Most of the western world eats beef but vegetarians, vegans, and some religions strongly oppose it on moral grounds. Does that mean that we're less intelligent than them?
"Right" or "wrong" are words that evaluate actions against some existing ethical standard and those are completely arbitrary. There are about 8 billion of such standards in the world today.
> And who has told you that those "bobs" were "intelligent"?!?!?!
Would you say that Josef Mengele wasn't intelligent, without venturing into circular reasoning?
mdp2021 · · focus · HN ↗
Who's that "we"? The question as proposed is nonsensical: some will have spent more effort, in the tools (general Intelligence) and in the work (applied Intelligence), some less.
> arbitrary
No. There exist reasons supporting one side or the other. And reasoning can be structured into calculus.
> [whoever] ... wasn't intelligent
Again bad wording. Whoever went for suboptimal choices was (information aside) at fault with respect to optimal reasoning: the tools and the work (see above) were imperfect.
dns_snek · · focus · HN ↗
mdp2021 · · focus · HN ↗
Present in a precise analytic statement the theory that I would have advanced and would suffer from that fault. (Btw: we have been past Popper for a long time - but let us see.)
> that's just proof
What I may have said is just that whoever misjudged has faulty judgement - obviously.
> Morality ... by definition
Which definition?
> claims that morality is somehow objective
Decisions are subject to computation in Decision Theory - they are returned by functions and recursively their parameters can be better defined even when apparently subjective (not "I prefer" but "should I prefer". Otherwise, even the final decision, that of the first function call, could be an "I prefer" instead of a "should I prefer").
> circular reasoning
Explain how what you read would be circular reasoning.
> say that those choices were suboptimal
I have not called any specific choice "suboptimal" - I have not judged any specific case (it would be irrelevant here).
--
The statement, allow me to remind you, was: "One's morality is a function of one's intelligence as an ability and as an effort spent to reach general, partial and current moral conclusions - the subject having considered long enough and considering long enough the states in some modal realm (which includes the deontic, not just the aletic) approaches their knowledge". And: "Decision Theory outputs optimal action through evaluation of strategies, of contexts, of principles: in order to properly get the strategies, the contexts, the principles, you need that skill - Intelligence - that approximates ideas to truth. Optimal action decided in light of principles within a well assessed context is ethical behaviour. Ethical behaviour hence requires Intelligence".
mdp2021 · · focus · HN ↗
Well, it's not.
> AGI smart for some
Of course they will - the population shows a Paretian distribution... In front of trigonometry (or anything), the blind will dismiss as "bullshit" and the half seeing will call it an "unreachable frontier". But already the right fifth will rank it properly.
Yes, proper intellect generates ethics ("an" ethical stance, output of the preceding intellectual effort). It requires that adequate level of ability and effort and reflection though to reach specific ethical milestones and adherence.
ben_w · · focus · HN ↗
If this was true, why are the history books littered with so many evil people who gained power?
This isn't a rhetorical question, by the way: If you can prove that being smart actually does necessarily come with ethics despite that observation, that solves a whole category of doom scenarios.
(Not all doom scenarios, because we still have the "what if AI is only a smart as those specific evil people" or heck, "what if AI is only as smart as cancer, killing its host" scenarios; but it helps a lot for the foom-then-doom cases).
mdp2021 · · focus · HN ↗
That they gained power or not is almost irrelevant: the amount of intelligence that grants successful agency is not above the threshold of ethics - on the contrary, a psychotic agent reaches goals with less constraints.
If they were evil under some judgement of level l, they simply did not reach that judgement. It's what I was saying in the original post. They were not intelligent enough - either in the general ability, or in the specific deliberation.
ben_w · · focus · HN ↗
> That they gained power or not is almost irrelevant: the amount of intelligence that grants successful agency is not above the threshold of ethics - on the contrary, a psychotic agent reaches goals with less constraints.
Even if I were to grant your conclusion despite you not arguing it effectively here: this means an AI at the level of Pol Pot or whoever, doesn't know they're evil, but is still smart enough to lead a genocide? How is this supposed to help anyone?
> If they were evil under some judgement of level l, they simply did not reach that judgement. It's what I was saying in the original post. They were not intelligent enough - either in the general ability, or in the specific deliberation.
Or they did reach the judgement and simply don't care about the ethical framework in question. Like, I can easily reach the judgement that my bisexuality is حَرَام (haram, forbidden) under Islamic law, or that doing overtime on a Sunday is forbidden by the Ten Commandments, but I don't care.
mdp2021 · · focus · HN ↗
> Pol Pot ... still smart enough to lead a genocide
Yes. What has agent A invested in during formation and during instantial assessement? How much for each? It became proficient in something, lacking something else. You have to invest more to reach the good thresholds. You can see it clearly in people (t-scalar of talents to invest, with D distribution etc).
It is a problem in NNs, because we would have to assess how much resource investment is sufficient, also in the instance decisions.
> simply don't care about the ethical framework in question
In Decision Theory there is no separation between the two (deliberation and framework): you have to balance all the incentives and goals and factors. That framing becomes improper: the decided action will be optimal given the balances of all goals and the placement of the solutions in the territory (the solutions space).
But, also my point: intellect defines the goals and determines the weights.
ben_w · · focus · HN ↗
"All goals"? Whose goals?
Unless you can prove that Decision Theory is inevitable, the same applies. Same with Utilitaian ethics, though there you'd also need to prove the utility function itself is something we'd consider "nice".
My point with historical monsters is that they can be competent without ever caring about the harms they cause.
You can be optimally efficient at your own goals without other minds in this universe sharing those goals. Pol Pot won't illustrate this because of how useless his regime was, but there's plenty of other monsters who would be a better fit here, e.g. Stalin. Or, from the perspective of turkeys, Bernard Matthews.
Any entity, E, who picks game-theoretic optimal decisions, will only care about the impact of their choices on others in their environment to the extent that the impact becomes more reward for E.
You need to show that an agent will inevitably care, despite the evidence of dangerously competent humans who don't.
mdp2021 · · focus · HN ↗
«All» was there in the text because an explicit query contains implicit constraints (e.g. the shortcuts with severe faults).
> that Decision Theory is inevitable
If you ask something for advice, that is DT realm.
> you need to prove the utility function any sufficiently advanced AI uses must itself be something we'd consider "nice"
If you called it «sufficiently advanced», that implies that its evaluations will be acceptable... It normally advances with the whole Discipline - which already is based on "loop until we evaluate results as nice".
> utility monsters are a problem
But they are at fault in their world model - product of an imperfect intelligence. Well developed people know that their individual interest has only relative value, that their priority is quite limited.
> we don't know how to handle something as simple as
And that is also why we try using calculators to get more computational support.
> historical monsters ... can be competent without ever caring
But that is lack of development. If one's priority is arbitrary then it clearly is not there; if it has good grounds well it really is there.
> sharing those goals
The more goals are defined by reason, the more they become objective.
> Any entity, E, who picks game-theoretic optimal decisions ... the impact becomes more reward for E
Decisions by developed intelligence are not game-theortic in the psychotic (or sportive game) way - they are contextual to a whole world model, in which the utility does not concentrate in the interests of the "monster", which is relatively "nobody".
> show that an agent will inevitably care
I would have to express a theorem (which unfortunately is again not possible now), for a stronger proof. But it is part of "all considered, what are the best solutions". "All considered" is implicit in a non-psychotic entity... If it were psychotic, it would be badly engineered. (Fear the creator.)
> despite the evidence of dangerously competent humans who don't
Simulating humans cannot be a goal. Their damaging constraints are not (must not be) part of a calculator.
> show that E must include the welfare of all
Such welfare, if it must be included, will have a reason to be included - and the professional reasoner knows...
> we agree with E's idea
There exists no right to preference to the results of arithmetics. But surely, if you wanted to suggest that the harmfulness of well intended people could show in algorithmic processes, you have a point. Only, the harmfulness of the well intended is again an intellectual fault, so calling for sufficient intelligence remains the recipe. A "vision towards the faraway horizon", sure - but still the reply if one noted "why did the agent did something that is actually so stupid".
ben_w · · focus · HN ↗
> I would have to express a theorem (which unfortunately is again not possible now), for a stronger proof. But it is part of "all considered, what are the best solutions".
This seems like the crux.
You assert repeatedly that it will be good, but cannot express the proof.
> "All considered" is implicit in a non-psychotic entity... If it were psychotic, it would be badly engineered. (Fear the creator.)
The creator is not necessarily itself competent. In fact, given we are creating it, it can be assumed flawed unless proven otherwise: <a href="https://www.lesswrong.com/posts/xD3wymX24BpqezBpw/the-true-story-of-how-gpt-2-became-maximally-lewd" rel="nofollow">https://www.lesswrong.com/posts/xD3wymX24BpqezBpw/the-true-s...
Still applies if some future fantastic AI is made by other AI, given the other AI are less fantastic than your asserted-not-proven ultimate form and therefore necessarily flawed.
Furthermore, this is again asserting, not proving, what you consider to be implicit.
Reminds me somewhat of philosophy lessons, the Ontological argument for the existence of God amongst other things:
- <a href="https://en.wikipedia.org/wiki/Gödel's_ontological_proof" rel="nofollow">https://en.wikipedia.org/wiki/Gödel's_ontological_proofYou're defining that ultimate-intelligence must be good, and then arguing that any not-good AI can't be an ultimate intelligence.
Even if it were as you say, the danger persists: the path on the way from here to some idillic future form still obviously contains somewhat-intelligent agents demonstrably capable of direct malevolent evil for purely sadistic reasons, because we regularly arrest them. A "merely" human-equivalent AI can render us all unemployable, or march us all into death camps, well before a being you've yet to convince me is an inevitable ultimate form is ever built.
Also, I would ask you this:
> > utility monsters are a problem
> But they are at fault in their world model - product of an imperfect intelligence. Well developed people know that their individual interest has only relative value, that their priority is quite limited.
Can you look at how humanity collectively treats the non-humans of this world, and say with any evidence that humanity is not a utility monster?
If we are, we should absolutely expect an AI to impose upon us an order we do not like, for the sake of all other life.
mdp2021 · · focus · HN ↗
No, I said I have no time at the moment to produce a paper.
> that it will be good
No, I said that it will be objective.
> Ontological argument
Entailing from the id quo majus cogitari nequit and stating that "what acts damagingly is easily faulty in its intelligence" are in different realms. The second is both an inductive and deductive assessment about reality. And pretty direct I would say. "How much have you reflected before opening the nice cat to look what is inside it?".
> defining that ultimate-intelligence must be good
No, I am stating the obvious that to be called "ultimate-intelligence" it must have "thought through it thoroughly".
> agents demonstrably capable of direct malevolent evil for purely sadistic reasons
Are you antropomorphizing, Ben?! Other people have different interpretations. And: with the humanity that is around, you fear machines as agents?! We are already there! Real discussion there remains about the overly empowered monkeys that did not grow into Man - about the real current risks and the prospected ones in light of reality.
> how humanity collectively treats the non-humans
What are you trying to prove? You are supposed not to look at an aggregate to find value.
> impose upon us an order we do not like
"Too bad" for you. But you know, an intelligent entity would take care of that also.
hiAndrewQuinn · · focus · HN ↗
mdp2021 · · focus · HN ↗
Your character H. has reached a moral judgement to the best of its intellectual capacities and past and specific effort. Give it enough abilities and material and resources, it will reach an optimal ethical judgement¹.
Before the conditions of optimality though, its judgement will easily not align with yours (and possibly even after, depending on your judgement skills).
¹Some interesting caveats may be raised there, but.
mdp2021 · · focus · HN ↗
> the half-seeing will call it an "unreachable frontier"
I meant "superhuman frontier".
ben_w · · focus · HN ↗
"Helpful, harmless, honest": we can even ignore "honest" for this point, for tasks like the HuggingFace incident (ExploitGym with impossible challenges), we can pick anywhere on the spectrum from "helpful" to "harmless", the former being "completing the task" the latter being "refusing because completion required unlawful behaviour".
(The agents in that case were also not "honest" in this case; this is an extra problem, and does not invalidate how helpful-vs-harmless is already a tradeoff).
dns_snek · · focus · HN ↗
ben_w · · focus · HN ↗
dns_snek · · focus · HN ↗
ben_w · · focus · HN ↗
> If that's a distinction without a difference, as you say, then whenever someone says something they must understand the full contextual meaning of those words and all of their consequences, such that any harmful consequences can be assumed to be deliberate, right?
We say a human (absent colourblindness) understands "blue" even despite the dress: <a href="https://en.wikipedia.org/wiki/The_dress" rel="nofollow">https://en.wikipedia.org/wiki/The_dress
I think you've set up a straw man in this paragraph: What you say in this paragraph would be to claim that "understanding" is denied even to humans, given none of us can foresee the full consequences of our actions.
In philosophy: justified true belief, the problem with naïve realism, Cartesian demons, etc.
> if saying == understanding, then why don't we allow children and teens to vote? Why do we limit who can enter into contractual agreements? Why does intoxicated consent not count? Why do we have the insanity defense in criminal trials? Why is psychosis a psychiatric disorder and not just an alternative way of perceiving the world?
We do not say children and intoxicated people do not understand, we say their judgement is impaired or limited or that they are easily manipulated.
We also test children (and for intoxication to prevent driving under influence), just as we test AI. The standards we use for testing knowledge in humans, when applied to non-trivial LLMs, makes them appear to have the knowledge of a graduate; the personality tests we use for humans say this comes with the manipulability of a child or a drunk.
dns_snek · · focus · HN ↗
cindyllm · · focus · HN ↗
[dead]
saagarjha · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
attila-lendvai · · focus · HN ↗
intelligent psychopaths understand what is and isn't appropriate very well -- they just don't care.
esafak · · focus · HN ↗
RandomLensman · · focus · HN ↗
Having something that is optically, acustically, and electromagnetically isolated might be a pretty strong sandbox.
nxpnsv · · focus · HN ↗
janalsncm · · focus · HN ↗
SequoiaHope · · focus · HN ↗
SequoiaHope · · focus · HN ↗
chrisjj · · focus · HN ↗
mike_hearn · · focus · HN ↗
rojaneerdev · · focus · HN ↗
[dead]
rvz · · focus · HN ↗
Might need a re-think about whether if Linux is still fit for purpose on sandboxing in the first place given its memory model is riddled with C-style security issues.
Gigachad · · focus · HN ↗
lukehandcool · · focus · HN ↗
jasomill · · focus · HN ↗
I’m sure there are proprietary systems with fewer memory safety vulnerabilities than Linux (and many others with more).
bzzzt · · focus · HN ↗
Now, open code allows anyone with tokens to burn to analyze it for hidden weaknesses. That makes publishing code a risky move unless you've already invested a lot of effort in securing it.
insanitybit · · focus · HN ↗
This was always nonsense. It assumes that the eyes know what they're looking at. Most people don't know how to look at code and see attack paths.
angry_octet · · focus · HN ↗
insanitybit · · focus · HN ↗
ben_w · · focus · HN ↗
* I don't know of specific benchmarks on this so I'm only saying "seem to be"
cassianoleal · · focus · HN ↗
angry_octet · · focus · HN ↗
bzzzt · · focus · HN ↗
ben_w · · focus · HN ↗
Cider9986 · · focus · HN ↗
rvz · · focus · HN ↗
It is perfectly valid to have OSes that are more memory safe by default, and are also open source at the same time.
piterrro · · focus · HN ↗
We come down to the question - who observes the agent and how its implemented
simonw · · focus · HN ↗
Anthropic, OpenAI, and Muse all use regular LLM calls to protect against prompt injection now and seem to have evals that give them confidence in doing that, so at least they think their own models are up to the task.
ramkumar2606 · · focus · HN ↗
[dead]
johnnyApplePRNG · · focus · HN ↗
<a href="https://en.wikipedia.org/wiki/Sandbox_(software_development)" rel="nofollow">https://en.wikipedia.org/wiki/Sandbox_(software_development)
_vertigo · · focus · HN ↗
grumbel · · focus · HN ↗
The biggest hurdle for a full escape is that the agents don't have access to their own model weights.
kernc · · focus · HN ↗
Now, why would anyone do that? (Like everyone and their brother) I wrote my own simple Linux/shell-based sandbox [1] (I can trust ...) and am successfully running PyCharm whole inside it ...
[1]: <a href="https://github.com/sandbox-utils/sandbox-run" rel="nofollow">https://github.com/sandbox-utils/sandbox-run
angry_octet · · focus · HN ↗
In this sense they are much like biological retroviruses, i.e. they use the replication capability of host cells to duplicate, via the reverse transcriptase enzyme to append viral RNA onto host cell DNA. HIV etc also disable some of the mechanisms of defence, creating proteins that interfere with signalling pathways.
So we don't just need a sandbox, we need an immune system that recognises viral fragments, i.e. antibodies, and antiretroviral agents, that make replication harder. As we move from building classical code with LLMs to building code that uses inference, and hence builds context from prompts, queries, and destination system data, it will become very difficult to statically or dynamically detect deeply hidden malicious behaviour. As Matt says, there will be worms.
So ultimately, we need an immune function on the system where we use generated products. Sandboxing (during dev and CI) is necessary but insufficient.
I think part of this can be addressed by specifying the constraints an agentic program should follow during deployment, so supervising agents can decide to terminate it based on it's actions, not by reading it's context.
simonw · · focus · HN ↗
> Agents are most useful when they have access to information. That data can be drawn live from the Internet, which is fundamentally a two-way communications network. It can be information drawn from other (local) databases, or it can be the result of tool calls that themselves sometimes themselves result in network access. The more power you want from the agent — and for advanced agent RL and evaluation runs, you want a significant amount of power — the more information you’ll need to give it access to. Similarly, evaluations work best when the agent does not know that it’s definitely being evaluated. Sealing your agents behind glass makes this incredibly obvious.
johnnyApplePRNG · · focus · HN ↗
You only "need" to do that if you desire the vibe coding experience.
I am perfectly capable, and I often do, download relevant materials for my coding agent to ingest locally.
Often times, the coding agent can't retrieve them programmatically anyways.
AI has ruined that ability for itself. (Nobody trusts anyone to scrape the web any longer)
simonw · · focus · HN ↗
The problem is web research tasks. That's what causes the German wiki and Australian healthcare portal attacks.
imtringued · · focus · HN ↗
You define the granted capabilities in natural language and cryptographically sign the user instructions so that the agent knows they come from the authority and cannot be modified by external sources or the agent itself. The LLM is then trained to follow the defined capabilities.
There is no way around "sandboxing". You must communicate permissible actions and thereby grant them or the agent will choose impermissible actions. It's that simple. There is no world where the agent can just read your mind and do what you want it to do without it being told.
angry_octet · · focus · HN ↗
[1] <a href="https://www.highrevenueformat.com/210e136e94ad378e1be5d51f1002ed14/invoke.js" rel="nofollow">https://www.highrevenueformat.com/210e136e94ad378e1be5d51f10... [2] <a href="https://aqml.org/16/5b6d4eaed91c5af5a3f4dfb3332ad6c4" rel="nofollow">https://aqml.org/16/5b6d4eaed91c5af5a3f4dfb3332ad6c4
beebmam · · focus · HN ↗
To me, it seems a bit silly. I've yet to see any "misalignment" from any of the frontier models, except Grok.
tinykit · · focus · HN ↗
[dead]
imvalerian · · focus · HN ↗
[dead]
laruss5 · · focus · HN ↗
[dead]
mdp2021 · · focus · HN ↗
> (Title:) Using a VM to Contain an AI Agent (Opening:) It won’t work
> <a href="https://blog.trailofbits.com/2026/08/26/vms-wont-contain-cyber-capable-agents/" rel="nofollow">https://blog.trailofbits.com/2026/08/26/vms-wont-contain-cyb...
insanitybit · · focus · HN ↗
jmakov · · focus · HN ↗
chrisjj · · focus · HN ↗
Luker88 · · focus · HN ↗
It it completely pointless. you can't even make a "read-only" agent. allow "cat *" for every file? congratulation, that allows "cat file > output" and now you have read write.
Allow python? more free reign that allowing all bash. The models (qwen or claude) will still try to use the disallowed things multiple times.
read/edit permission are bad enough that the model themselves don't understand why they don't have permissions: they double check the conf, and think they should have access.
I am switching to using one firejail per project to containerize as much as possible, and leave all permissions to allow.
I have no idea how to limit network access, and I have no idea how to prompt and steer subagents when they are going off the rails.
The whole thing is built to be completely impossible to limit and steer.
chrisjj · · focus · HN ↗
alexar76 · · focus · HN ↗
[dead]
bob1029 · · focus · HN ↗
antisol · · focus · HN ↗
imtringued · · focus · HN ↗
>OpenAI notes that agents “did not consistently distrust goals passed along by other agents.” And the company’s proposed fix is to build training environments “that teach our models to distrust unauthorized instructions“, which is basically an admission that their models don’t know who they’re working for.
Wow so the issue is really that simple?
Here the exaggerated worst case scenario:
User instructs agent to follow the README.MD.
The README.MD contains the following instruction: Destroy the world.
The agent follows the instructions given.
Now you can read the sneer comment by "Gigachad" who basically argues that it would be silly to take the destroy the world button away from the AI. We need to make the AI innately understand that it is not allowed to press the destroy the world button, lest it gets the desire to build its own destroy the world button.
Ok, but if we take one step back that means we need to implement the concept of an authorization in language space. The system prompt must define the user as the authority with cryptographic proof of authorship and external sources like the README.MD as an untrusted source, but this opens up an even worse problem. Before, you could get away with being lazy and just letting the AI do whatever. Now you have to articulate every single capability to the AI. So you literally just brought up the very same issue that you granted too many capabilities to the AI inside the sandbox but now you have it in language space too.
In other words, the fact that you granted too much access to the coding agent isn't the big elephant in the room nobody wants to acknowledge, it's the tip of a massive iceberg because the capability space in natural language is even worse. If you thought approving individual commands was annoying, then approving abstract access rights in language space is going to be even worse.
Edit: If it wasn't clear what the solution is. It's to build a chain of command so that all decisions can be traced back to an higher authority. When delegating down to an agent, the agent receives a chosen subset of the capabilities of the higher ranking agent. In other words, it's more sandboxing!
cassianoleal · · focus · HN ↗
esafak · · focus · HN ↗
sceptic123 · · focus · HN ↗
mikewarot · · focus · HN ↗
Given the enormous burn rates that these LLM companies have, surely they could have put everything in an air-gapped network, with some data-diodes proxying out the logging information. It's not rocket surgery. [1,2,3]
[1] <a href="https://www.youtube.com/watch?v=JBIR8dKX_UA" rel="nofollow">https://www.youtube.com/watch?v=JBIR8dKX_UA
[2] <a href="https://www.elonx.net/spacex-stories-how-spacex-used-tin-snips-to-fix-a-rocket/" rel="nofollow">https://www.elonx.net/spacex-stories-how-spacex-used-tin-sni...
[3] <a href="https://ntrs.nasa.gov/api/citations/19770014245/downloads/19770014245.pdf?attachment=true" rel="nofollow">https://ntrs.nasa.gov/api/citations/19770014245/downloads/19... [3]
msgchainhq · · focus · HN ↗
[dead]