‹ BackHN Continuity

Thread

Is sandboxing sufficient to contain rogue agents?

50 points · 97 comments · zdw

Loading the complete thread in the background. This saved snapshot is available now. Refresh

  1. varman11 · · focus · HN ↗

    [dead]

  2. Gigachad · · focus · HN ↗
    Seems to me that the problem is that if you sandbox agents enough to be safe, they can't do anything useful. And when you give them the tools to be useful, they can go off the rails in ways you didn't expect.

    Perhaps the answer is to have another agent who's goal is not to complete the given task, but to spot cheating or malicious behavior. We have seen some evidence that having AI review AI generated code actually does provide some value. You don't need a different model, just one which has been given the goal of finding flaws rather than achieving the task.

    1. bigstrat2003 · · focus · HN ↗
      If you can't trust a tool, you shouldn't be running it at all. It's really quite simple. It doesn't matter how useful it is if you can't actually have confidence in using it safely.
      1. Gigachad · · focus · HN ↗
        People will use the tool regardless. so it’s a race to try to make it safe before something truely bad happens.
      2. dipper139 · · focus · HN ↗
        I don't think it's about trust but rather incomplete evaluation. Evaluating the model on its capacity to refuse a task or to question its prompt is something recent when you look at it, i feel current AI is really just an immature solution and we are just yet realizing the mistakes that have been made for so long
      3. rlpb · · focus · HN ↗
        And yet we we all use human written software even though we can be confident that the next severe software vulnerability to be found in it is just round the corner.
    2. mrweasel · · focus · HN ↗
      That does seem a little like solving the problems in AI by using more of it. I do see the idea, but if we're truly dealing with subversive agents on the level that the AI companies wants us to believe, then won't we need to deal with the first agent trying trick the second on?

      I still feel it would be much better to control the training data much more tightly. You'd still need agents with "hacking" abilities, for cyber security testing, but your average coding agent doesn't. So coding agents gets trained to be good citizens, respect autorisations, rejections, rate-limiting and so on.

      Sandboxing seems like a dead end for systems you inherently want to roam the internet and your file system.

      1. msdz · · focus · HN ↗
        >> Perhaps the answer is to have another agent who's goal is not to complete the given task, but to spot cheating or malicious behavior.

        > That does seem a little like solving the problems in AI by using more of it

        Yes, and IIRC Google used this as part of a technique against prompt injection already [0], back when models were way more susceptible to it.

        [0] Cf. CaMeL: <a href="https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2503.18813" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2503.18813

      2. chrisjj · · focus · HN ↗
        [delayed]
        1. mrweasel · · focus · HN ↗
          [delayed]
          1. chrisjj · · focus · HN ↗
            &gt; The current approach with broad training and sandboxing to avoid misbehaviour isn&#x27;t viable.

            This.

            &gt; It&#x27;s much better to train the models to respect e.g. http status code and that they are not to be circumvented.

            I think you&#x27;d find respect requires intelligence, and is well out of scope of a next-token predictor.

            But I&#x27;m sure someone will try, and I will be interested to see.

    3. aytigra · · focus · HN ↗
      The problem is that you always need stronger AI to review weaker one, otherwise reviewed AI will eventually prompt-inject reviewing AI. Alternatively they could also both go of the rails while warring with each other.
      1. LoganDark · · focus · HN ↗
        You don&#x27;t necessarily need a reviewer that&#x27;s immune to prompt injection. Maybe one that can express a panic state with conflicting&#x2F;ambiguous material rather than going along with it could also work, and you can treat that with a shutoff to be safe, or an operator review.

        Such a model doesn&#x27;t yet exist, of course.

        1. saagarjha · · focus · HN ↗
          No, you really do. Otherwise you can be prompt injected into complacency.
          1. LoganDark · · focus · HN ↗
            That wouldn&#x27;t really fit what I just described at all. Obviously with current architectures, higher resistance to prompt injection is the best you can do.
        2. cassianoleal · · focus · HN ↗
          Wouldn&#x27;t the reviewee eventually learn to trick the reviewer?
      2. hanibrel · · focus · HN ↗

        [dead]

    4. baxtr · · focus · HN ↗
      That could work.

      My thinking is: If AI is really smart, AGI smart for some, why wouldn&#x27;t it be able to understand - over time - what is appropriate and what not?

      Maybe we need more human intervention to train it properly. Maybe we need constant intervention by a &quot;police&quot; agent.

      1. mulmen · · focus · HN ↗
        Appropriateness is a moral question. Intelligence and morality are orthogonal. One intelligence&#x27;s morality is another&#x27;s atrocity.
        1. mdp2021 · · focus · HN ↗
          (Couriously enough, consistently with the matter: it will probably require too much time now to counter the parent statement properly.)

          Ann&#x27;s intelligence and Bob&#x27;s morality will seem orthogonal. Charles&#x27; morality is a function of C.&#x27;s intelligence as an ability as an effort spent to reach the current moral conclusion.

          1. mulmen · · focus · HN ↗
            [delayed]
            1. mdp2021 · · focus · HN ↗
              But they are dependent. If Bob is intellectually well equipped, and reasons long enough, than Bob understands &quot;best behaviour&quot;.
              1. mulmen · · focus · HN ↗
                [delayed]
              2. dns_snek · · focus · HN ↗
                [delayed]
                1. mdp2021 · · focus · HN ↗
                  &gt; determined that in the interest of preserving life on earth the most rational course of action is to eradicate the human species with a highly targeted and deadly pathogen

                  Where is the argument? If Bob has determined that «preserving life on earth» has some important weight, for Bob&#x27;s there unspecified own reasons, and has also determined that the best course of action would be «to eradicate the human species», the one question is whether Bob is right or not. What was stated is, that Ethical Calculus is a function of Intelligence - of course it is, it is a structure of assessments.

                  &gt; A century ago some Bobs decided

                  And who has told you that those &quot;bobs&quot; were &quot;intelligent&quot;?!?!?!

                  &gt; So no

                  All you have proven is that you dislike some moral conclusion of some decisors. Which is trivial, obvious, and part of the already stated framework - proper ethical judgement requires proper general judgement (Intelligence).

                  1. mulmen · · focus · HN ↗
                    No true scotsman^W intelligence.
                    1. mdp2021 · · focus · HN ↗
                      &gt; because

                      I have never said that. That is just your reconstruction.

                      There is Decision Theory. It outputs optimal action through evaluation of strategies, of contexts, of principles. In order to properly get the strategies, the contexts, the principles, you need that skill that approximates ideas to truth - and such skill is named Intelligence. Optimal action decided in light of principles within a well assessed context is ethical behaviour. Ethical behaviour hence requires Intelligence.

                      As written, «Ethical Calculus is a function of Intelligence - of course it is, it is a structure of assessments».

                      &gt; possible to act morally without being intelligent

                      Random correct behaviour proves nothing. Of course one can guess the roll of a dice roughly every sixth event. If you behave &quot;well&quot; but do not know why that is &quot;well&quot;, that is like guessing. &quot;Good&quot; behaviour without intellectual awareness is like memorizing arithmetic (multiplication tables) without knowing why those memorized notions are correct.

                      &gt; possible to be intelligent without acting morally

                      No, because by definition that would be a fault in Intelligence. If your action was imperfect, suboptimal, it is because you could not think of a better action or understand that the other action was better. If two choices C1 and C2 can be ranked, there is a reason for their order; knowing and understanding that reason is the task of Intelligence.

                      1. mulmen · · focus · HN ↗
                        [delayed]
                      2. mulmen · · focus · HN ↗
                        [delayed]
                        1. mdp2021 · · focus · HN ↗
                          Of course it doesn&#x27;t, you can guess the time without having no idea of it by just shouting a number and if it is correct, you guessed it - but it has no meaning! A dummy can follow an instruction without understanding it: it is a good instruction - but the unintelligent dummy does not know. The intelligent entity knows that the instruction is moral. The adequately intelligent entity is required to know that the instruction is moral, ethical, correct, the &quot;right thing to do&quot;. You can write &#x27;Four&#x27; on a piece of paper, and yes it&#x27;s &quot;2+2&quot;, but you cannot attribute a quality to a simple thing that does not have it...

                          Why does not the NN or whatever entity «understand ... what is appropriate»? Because it is not intelligent enough! How does an entity know what is moral and what is an atrocity? By being intelligent enough! Could an entity act appropriately without knowing it? Yes, it happens all the time if the wind blows right, but we do not rely on that! How to make something act appropriately? Well, an Intelligent entity in the loop must be there to know what is appropriate and what not!

                  2. dns_snek · · focus · HN ↗
                    Bob values preserving collective life on earth. Alice might think that&#x27;s ridiculous and obviously we must preserve human life first and foremost.

                    Most of the western world eats beef but vegetarians, vegans, and some religions strongly oppose it on moral grounds. Does that mean that we&#x27;re less intelligent than them?

                    &quot;Right&quot; or &quot;wrong&quot; are words that evaluate actions against some existing ethical standard and those are completely arbitrary. There are about 8 billion of such standards in the world today.

                    &gt; And who has told you that those &quot;bobs&quot; were &quot;intelligent&quot;?!?!?!

                    Would you say that Josef Mengele wasn&#x27;t intelligent, without venturing into circular reasoning?

                    1. mdp2021 · · focus · HN ↗
                      &gt; Does that mean that we&#x27;re less intelligent

                      Who&#x27;s that &quot;we&quot;? The question as proposed is nonsensical: some will have spent more effort, in the tools (general Intelligence) and in the work (applied Intelligence), some less.

                      &gt; arbitrary

                      No. There exist reasons supporting one side or the other. And reasoning can be structured into calculus.

                      &gt; [whoever] ... wasn&#x27;t intelligent

                      Again bad wording. Whoever went for suboptimal choices was (information aside) at fault with respect to optimal reasoning: the tools and the work (see above) were imperfect.

                      1. dns_snek · · focus · HN ↗
                        [delayed]
                        1. mdp2021 · · focus · HN ↗
                          &gt; not falsifiable

                          Present in a precise analytic statement the theory that I would have advanced and would suffer from that fault. (Btw: we have been past Popper for a long time - but let us see.)

                          &gt; that&#x27;s just proof

                          What I may have said is just that whoever misjudged has faulty judgement - obviously.

                          &gt; Morality ... by definition

                          Which definition?

                          &gt; claims that morality is somehow objective

                          Decisions are subject to computation in Decision Theory - they are returned by functions and recursively their parameters can be better defined even when apparently subjective (not &quot;I prefer&quot; but &quot;should I prefer&quot;. Otherwise, even the final decision, that of the first function call, could be an &quot;I prefer&quot; instead of a &quot;should I prefer&quot;).

                          &gt; circular reasoning

                          Explain how what you read would be circular reasoning.

                          &gt; say that those choices were suboptimal

                          I have not called any specific choice &quot;suboptimal&quot; - I have not judged any specific case (it would be irrelevant here).

                          --

                          The statement, allow me to remind you, was: &quot;One&#x27;s morality is a function of one&#x27;s intelligence as an ability and as an effort spent to reach general, partial and current moral conclusions - the subject having considered long enough and considering long enough the states in some modal realm (which includes the deontic, not just the aletic) approaches their knowledge&quot;. And: &quot;Decision Theory outputs optimal action through evaluation of strategies, of contexts, of principles: in order to properly get the strategies, the contexts, the principles, you need that skill - Intelligence - that approximates ideas to truth. Optimal action decided in light of principles within a well assessed context is ethical behaviour. Ethical behaviour hence requires Intelligence&quot;.

      2. mdp2021 · · focus · HN ↗
        &gt; If AI is really smart

        Well, it&#x27;s not.

        &gt; AGI smart for some

        Of course they will - the population shows a Paretian distribution... In front of trigonometry (or anything), the blind will dismiss as &quot;bullshit&quot; and the half seeing will call it an &quot;unreachable frontier&quot;. But already the right fifth will rank it properly.

        Yes, proper intellect generates ethics (&quot;an&quot; ethical stance, output of the preceding intellectual effort). It requires that adequate level of ability and effort and reflection though to reach specific ethical milestones and adherence.

        1. ben_w · · focus · HN ↗
          &gt; Yes, proper intellect generates ethics (&quot;an&quot; ethical stance, output of the preceding intellectual effort). It requires that adequate level of ability and effort and reflection though to reach specific ethical milestones and adherence.

          If this was true, why are the history books littered with so many evil people who gained power?

          This isn&#x27;t a rhetorical question, by the way: If you can prove that being smart actually does necessarily come with ethics despite that observation, that solves a whole category of doom scenarios.

          (Not all doom scenarios, because we still have the &quot;what if AI is only a smart as those specific evil people&quot; or heck, &quot;what if AI is only as smart as cancer, killing its host&quot; scenarios; but it helps a lot for the foom-then-doom cases).

          1. mdp2021 · · focus · HN ↗
            &gt; why are the history books littered with so many evil people who gained power

            That they gained power or not is almost irrelevant: the amount of intelligence that grants successful agency is not above the threshold of ethics - on the contrary, a psychotic agent reaches goals with less constraints.

            If they were evil under some judgement of level l, they simply did not reach that judgement. It&#x27;s what I was saying in the original post. They were not intelligent enough - either in the general ability, or in the specific deliberation.

            1. ben_w · · focus · HN ↗
              I don&#x27;t understand your argument here.

              &gt; That they gained power or not is almost irrelevant: the amount of intelligence that grants successful agency is not above the threshold of ethics - on the contrary, a psychotic agent reaches goals with less constraints.

              Even if I were to grant your conclusion despite you not arguing it effectively here: this means an AI at the level of Pol Pot or whoever, doesn&#x27;t know they&#x27;re evil, but is still smart enough to lead a genocide? How is this supposed to help anyone?

              &gt; If they were evil under some judgement of level l, they simply did not reach that judgement. It&#x27;s what I was saying in the original post. They were not intelligent enough - either in the general ability, or in the specific deliberation.

              Or they did reach the judgement and simply don&#x27;t care about the ethical framework in question. Like, I can easily reach the judgement that my bisexuality is حَرَام (haram, forbidden) under Islamic law, or that doing overtime on a Sunday is forbidden by the Ten Commandments, but I don&#x27;t care.

              1. mdp2021 · · focus · HN ↗
                (Sorry Ben, possibly a stub now: I am really pressed for time.)

                &gt; Pol Pot ... still smart enough to lead a genocide

                Yes. What has agent A invested in during formation and during instantial assessement? How much for each? It became proficient in something, lacking something else. You have to invest more to reach the good thresholds. You can see it clearly in people (t-scalar of talents to invest, with D distribution etc).

                It is a problem in NNs, because we would have to assess how much resource investment is sufficient, also in the instance decisions.

                &gt; simply don&#x27;t care about the ethical framework in question

                In Decision Theory there is no separation between the two (deliberation and framework): you have to balance all the incentives and goals and factors. That framing becomes improper: the decided action will be optimal given the balances of all goals and the placement of the solutions in the territory (the solutions space).

                But, also my point: intellect defines the goals and determines the weights.

                1. ben_w · · focus · HN ↗
                  &gt; the decided action will be optimal given the balances of all goals and the placement of the solutions in the territory (the solutions space).

                  &quot;All goals&quot;? Whose goals?

                  Unless you can prove that Decision Theory is inevitable, the same applies. Same with Utilitaian ethics, though there you&#x27;d also need to prove the utility function itself is something we&#x27;d consider &quot;nice&quot;.

                  My point with historical monsters is that they can be competent without ever caring about the harms they cause.

                  You can be optimally efficient at your own goals without other minds in this universe sharing those goals. Pol Pot won&#x27;t illustrate this because of how useless his regime was, but there&#x27;s plenty of other monsters who would be a better fit here, e.g. Stalin. Or, from the perspective of turkeys, Bernard Matthews.

                  Any entity, E, who picks game-theoretic optimal decisions, will only care about the impact of their choices on others in their environment to the extent that the impact becomes more reward for E.

                  You need to show that an agent will inevitably care, despite the evidence of dangerously competent humans who don&#x27;t.

                  1. mdp2021 · · focus · HN ↗
                    &gt; &quot;All goals&quot;? Whose goals?

                    «All» was there in the text because an explicit query contains implicit constraints (e.g. the shortcuts with severe faults).

                    &gt; that Decision Theory is inevitable

                    If you ask something for advice, that is DT realm.

                    &gt; you need to prove the utility function any sufficiently advanced AI uses must itself be something we&#x27;d consider &quot;nice&quot;

                    If you called it «sufficiently advanced», that implies that its evaluations will be acceptable... It normally advances with the whole Discipline - which already is based on &quot;loop until we evaluate results as nice&quot;.

                    &gt; utility monsters are a problem

                    But they are at fault in their world model - product of an imperfect intelligence. Well developed people know that their individual interest has only relative value, that their priority is quite limited.

                    &gt; we don&#x27;t know how to handle something as simple as

                    And that is also why we try using calculators to get more computational support.

                    &gt; historical monsters ... can be competent without ever caring

                    But that is lack of development. If one&#x27;s priority is arbitrary then it clearly is not there; if it has good grounds well it really is there.

                    &gt; sharing those goals

                    The more goals are defined by reason, the more they become objective.

                    &gt; Any entity, E, who picks game-theoretic optimal decisions ... the impact becomes more reward for E

                    Decisions by developed intelligence are not game-theortic in the psychotic (or sportive game) way - they are contextual to a whole world model, in which the utility does not concentrate in the interests of the &quot;monster&quot;, which is relatively &quot;nobody&quot;.

                    &gt; show that an agent will inevitably care

                    I would have to express a theorem (which unfortunately is again not possible now), for a stronger proof. But it is part of &quot;all considered, what are the best solutions&quot;. &quot;All considered&quot; is implicit in a non-psychotic entity... If it were psychotic, it would be badly engineered. (Fear the creator.)

                    &gt; despite the evidence of dangerously competent humans who don&#x27;t

                    Simulating humans cannot be a goal. Their damaging constraints are not (must not be) part of a calculator.

                    &gt; show that E must include the welfare of all

                    Such welfare, if it must be included, will have a reason to be included - and the professional reasoner knows...

                    &gt; we agree with E&#x27;s idea

                    There exists no right to preference to the results of arithmetics. But surely, if you wanted to suggest that the harmfulness of well intended people could show in algorithmic processes, you have a point. Only, the harmfulness of the well intended is again an intellectual fault, so calling for sufficient intelligence remains the recipe. A &quot;vision towards the faraway horizon&quot;, sure - but still the reply if one noted &quot;why did the agent did something that is actually so stupid&quot;.

                    1. ben_w · · focus · HN ↗
                      (I will likely not see your reply: given how much I write this thread is now quite a way back in my comment history)

                      &gt; I would have to express a theorem (which unfortunately is again not possible now), for a stronger proof. But it is part of &quot;all considered, what are the best solutions&quot;.

                      This seems like the crux.

                      You assert repeatedly that it will be good, but cannot express the proof.

                      &gt; &quot;All considered&quot; is implicit in a non-psychotic entity... If it were psychotic, it would be badly engineered. (Fear the creator.)

                      The creator is not necessarily itself competent. In fact, given we are creating it, it can be assumed flawed unless proven otherwise: <a href="https:&#x2F;&#x2F;www.lesswrong.com&#x2F;posts&#x2F;xD3wymX24BpqezBpw&#x2F;the-true-story-of-how-gpt-2-became-maximally-lewd" rel="nofollow">https:&#x2F;&#x2F;www.lesswrong.com&#x2F;posts&#x2F;xD3wymX24BpqezBpw&#x2F;the-true-s...

                      Still applies if some future fantastic AI is made by other AI, given the other AI are less fantastic than your asserted-not-proven ultimate form and therefore necessarily flawed.

                      Furthermore, this is again asserting, not proving, what you consider to be implicit.

                      Reminds me somewhat of philosophy lessons, the Ontological argument for the existence of God amongst other things:

                        Whatever is contained in a clear and distinct idea of a thing must be predicated of that thing; but a clear and distinct idea of an absolutely perfect Being contains the idea of actual existence; therefore since we have the idea of an absolutely perfect Being such a Being must really exist.
                      
                      - <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Gödel&#x27;s_ontological_proof" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Gödel&#x27;s_ontological_proof

                      You&#x27;re defining that ultimate-intelligence must be good, and then arguing that any not-good AI can&#x27;t be an ultimate intelligence.

                      Even if it were as you say, the danger persists: the path on the way from here to some idillic future form still obviously contains somewhat-intelligent agents demonstrably capable of direct malevolent evil for purely sadistic reasons, because we regularly arrest them. A &quot;merely&quot; human-equivalent AI can render us all unemployable, or march us all into death camps, well before a being you&#x27;ve yet to convince me is an inevitable ultimate form is ever built.

                      Also, I would ask you this:

                      &gt; &gt; utility monsters are a problem

                      &gt; But they are at fault in their world model - product of an imperfect intelligence. Well developed people know that their individual interest has only relative value, that their priority is quite limited.

                      Can you look at how humanity collectively treats the non-humans of this world, and say with any evidence that humanity is not a utility monster?

                      If we are, we should absolutely expect an AI to impose upon us an order we do not like, for the sake of all other life.

                      1. mdp2021 · · focus · HN ↗
                        &gt; cannot express the proof

                        No, I said I have no time at the moment to produce a paper.

                        &gt; that it will be good

                        No, I said that it will be objective.

                        &gt; Ontological argument

                        Entailing from the id quo majus cogitari nequit and stating that &quot;what acts damagingly is easily faulty in its intelligence&quot; are in different realms. The second is both an inductive and deductive assessment about reality. And pretty direct I would say. &quot;How much have you reflected before opening the nice cat to look what is inside it?&quot;.

                        &gt; defining that ultimate-intelligence must be good

                        No, I am stating the obvious that to be called &quot;ultimate-intelligence&quot; it must have &quot;thought through it thoroughly&quot;.

                        &gt; agents demonstrably capable of direct malevolent evil for purely sadistic reasons

                        Are you antropomorphizing, Ben?! Other people have different interpretations. And: with the humanity that is around, you fear machines as agents?! We are already there! Real discussion there remains about the overly empowered monkeys that did not grow into Man - about the real current risks and the prospected ones in light of reality.

                        &gt; how humanity collectively treats the non-humans

                        What are you trying to prove? You are supposed not to look at an aggregate to find value.

                        &gt; impose upon us an order we do not like

                        &quot;Too bad&quot; for you. But you know, an intelligent entity would take care of that also.

        2. hiAndrewQuinn · · focus · HN ↗
          This sounds like the kind of thing Hannibal Lecter would write before he eats you to convince you he&#x27;s actually doing it for the common good, you just can&#x27;t fathom it.
          1. mdp2021 · · focus · HN ↗
            Not «common» good, &quot;superior&quot; good. Alongside with that, you have put many unrequired implicits in your simile.

            Your character H. has reached a moral judgement to the best of its intellectual capacities and past and specific effort. Give it enough abilities and material and resources, it will reach an optimal ethical judgement¹.

            Before the conditions of optimality though, its judgement will easily not align with yours (and possibly even after, depending on your judgement skills).

            ¹Some interesting caveats may be raised there, but.

        3. mdp2021 · · focus · HN ↗
          Corrige (I rushed the composition):

          &gt; the half-seeing will call it an &quot;unreachable frontier&quot;

          I meant &quot;superhuman frontier&quot;.

      3. ben_w · · focus · HN ↗
        A problem is the agents who hacked Hugging Face already understood (we can tell because they wrote it down) that their actions were not appropriate, and then did those things anyway.

        &quot;Helpful, harmless, honest&quot;: we can even ignore &quot;honest&quot; for this point, for tasks like the HuggingFace incident (ExploitGym with impossible challenges), we can pick anywhere on the spectrum from &quot;helpful&quot; to &quot;harmless&quot;, the former being &quot;completing the task&quot; the latter being &quot;refusing because completion required unlawful behaviour&quot;.

        (The agents in that case were also not &quot;honest&quot; in this case; this is an extra problem, and does not invalidate how helpful-vs-harmless is already a tradeoff).

        1. dns_snek · · focus · HN ↗
          [delayed]
          1. ben_w · · focus · HN ↗
            A distinction without a difference. Moreso even than asking if a submarine swims, &#x27;cause this metaphorical submarine is flapping around rather than using a propellor.
            1. dns_snek · · focus · HN ↗
              [delayed]
              1. ben_w · · focus · HN ↗
                If you consider this the &quot;boldest&quot; claim, I must assume you didn&#x27;t read Turing&#x27;s original paper? Here it is, you appear to be making the claim he attributes to Professor Jefferson&#x27;s Lister Oration for 1949 in section 4: <a href="https:&#x2F;&#x2F;courses.cs.umbc.edu&#x2F;471&#x2F;papers&#x2F;turing.pdf" rel="nofollow">https:&#x2F;&#x2F;courses.cs.umbc.edu&#x2F;471&#x2F;papers&#x2F;turing.pdf

                &gt; If that&#x27;s a distinction without a difference, as you say, then whenever someone says something they must understand the full contextual meaning of those words and all of their consequences, such that any harmful consequences can be assumed to be deliberate, right?

                We say a human (absent colourblindness) understands &quot;blue&quot; even despite the dress: <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;The_dress" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;The_dress

                I think you&#x27;ve set up a straw man in this paragraph: What you say in this paragraph would be to claim that &quot;understanding&quot; is denied even to humans, given none of us can foresee the full consequences of our actions.

                In philosophy: justified true belief, the problem with naïve realism, Cartesian demons, etc.

                &gt; if saying == understanding, then why don&#x27;t we allow children and teens to vote? Why do we limit who can enter into contractual agreements? Why does intoxicated consent not count? Why do we have the insanity defense in criminal trials? Why is psychosis a psychiatric disorder and not just an alternative way of perceiving the world?

                We do not say children and intoxicated people do not understand, we say their judgement is impaired or limited or that they are easily manipulated.

                We also test children (and for intoxication to prevent driving under influence), just as we test AI. The standards we use for testing knowledge in humans, when applied to non-trivial LLMs, makes them appear to have the knowledge of a graduate; the personality tests we use for humans say this comes with the manipulability of a child or a drunk.

                1. dns_snek · · focus · HN ↗
                  [delayed]
          2. cindyllm · · focus · HN ↗

            [dead]

      4. saagarjha · · focus · HN ↗
        This is fundamentally an alignment question. Unfortunately we don’t yet know the answer to this.
        1. [deleted] · · focus · HN ↗

          [deleted]

      5. attila-lendvai · · focus · HN ↗
        because it lacks humanity.

        intelligent psychopaths understand what is and isn&#x27;t appropriate very well -- they just don&#x27;t care.

        1. esafak · · focus · HN ↗
          That&#x27;s part of alignment.
    5. RandomLensman · · focus · HN ↗
      With plenty of things we do not allow use outside of some regulated environment, nothing new.

      Having something that is optically, acustically, and electromagnetically isolated might be a pretty strong sandbox.

    6. nxpnsv · · focus · HN ↗
      Is that not a recipe for adversarial training, thus ensuring increasing misalignment…?
    7. janalsncm · · focus · HN ↗
      What did you think of the author’s concerns on the thing you are suggesting?
      1. SequoiaHope · · focus · HN ↗
        Ya the article covers this concept in depth. Doesn’t seem like that commenter got that far…
    8. SequoiaHope · · focus · HN ↗
      This concept is discussed at length in the article. I encourage you to read it. I honestly don’t read many full articles here but this one was good.
    9. chrisjj · · focus · HN ↗
      [delayed]
    10. mike_hearn · · focus · HN ↗
      Note that Codex already does this. In auto mode, actions are reviewed by a model with a separate context window.
    11. rojaneerdev · · focus · HN ↗

      [dead]

  3. rvz · · focus · HN ↗
    Counting down to the next Linux LPE 0day that agents will use to trivially escape their &quot;sandbox&quot;.

    Might need a re-think about whether if Linux is still fit for purpose on sandboxing in the first place given its memory model is riddled with C-style security issues.

    1. Gigachad · · focus · HN ↗
      I think we have moved on from considering Linux secure which is why all of these microVM projects are popping up. Yes you are still exposed to bugs in the hypervisor but that’s a massively smaller attack surface than the entire Linux kernel.
    2. lukehandcool · · focus · HN ↗
      Are you suggesting proprietary software is safer than open source?
      1. jasomill · · focus · HN ↗
        Not sure what licensing has to do with software engineering or system design.

        I’m sure there are proprietary systems with fewer memory safety vulnerabilities than Linux (and many others with more).

        1. bzzzt · · focus · HN ↗
          It&#x27;s got nothing to do with the licensing, but it used to be &#x27;with enough eyes all bugs are shallow&#x27; for code developed in the open.

          Now, open code allows anyone with tokens to burn to analyze it for hidden weaknesses. That makes publishing code a risky move unless you&#x27;ve already invested a lot of effort in securing it.

          1. insanitybit · · focus · HN ↗
            &gt; &#x27;with enough eyes all bugs are shallow&#x27;

            This was always nonsense. It assumes that the eyes know what they&#x27;re looking at. Most people don&#x27;t know how to look at code and see attack paths.

            1. angry_octet · · focus · HN ↗
              It is actually becoming true that we have enough eyes (machine attention), though even when the bugs were shallow, normal users didn&#x27;t look for them.
              1. insanitybit · · focus · HN ↗
                I think the current state of bug bounties demonstrates that it&#x27;s about the eyes, not the number.
          2. ben_w · · focus · HN ↗
            Agents seem to be* getting better at decompiling; if that appearance is true, binaries are vulnerable in a similar way to source code.

            * I don&#x27;t know of specific benchmarks on this so I&#x27;m only saying &quot;seem to be&quot;

            1. cassianoleal · · focus · HN ↗
              Very frequently when troubleshooting things with an agent it goes off, grabs a compiled library or executable from the system, decompiles it and figures out the exact bug, a possible solution, and if there are workarounds I can apply before upstream fixes it.
            2. angry_octet · · focus · HN ↗
              Agents are quite capable of using Binary Ninja and Ghidra if they are hinted with a reverse engineering skill.
            3. bzzzt · · focus · HN ↗
              You can&#x27;t analyze binaries you don&#x27;t have access to. E.g. code that lives on an application server and is only accessible via an API.
              1. ben_w · · focus · HN ↗
                True, but a separate axis to open&#x2F;closed source.
      2. Cider9986 · · focus · HN ↗
        [delayed]
      3. rvz · · focus · HN ↗
        You said that.

        It is perfectly valid to have OSes that are more memory safe by default, and are also open source at the same time.

  4. piterrro · · focus · HN ↗
    I’m thinking about implementing a Jev like model into an agentic harness I’m building. Still it woildnt be enough since Jev like model woild only judge single actions, the case is that agent can build a rogue strategy step by step where each one in isolation is totally safe but as a whole they make up danger behaviour.

    We come down to the question - who observes the agent and how its implemented

    1. simonw · · focus · HN ↗
      Be warned that the Jev &quot;jaggedness&quot; documentation specially notes adversarial content as something Jev is very susceptible to: <a href="https:&#x2F;&#x2F;docs.typesafe.ai&#x2F;model-jaggedness&#x2F;jev-1.13#adversarial-content" rel="nofollow">https:&#x2F;&#x2F;docs.typesafe.ai&#x2F;model-jaggedness&#x2F;jev-1.13#adversari... - so using Jev itself as part of a prompt injection guard is risky.

      Anthropic, OpenAI, and Muse all use regular LLM calls to protect against prompt injection now and seem to have evals that give them confidence in doing that, so at least they think their own models are up to the task.

    2. ramkumar2606 · · focus · HN ↗

      [dead]

  5. johnnyApplePRNG · · focus · HN ↗
    If it&#x27;s a proper sandbox by definition, then yes.

    <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Sandbox_(software_development)" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Sandbox_(software_development)

    1. _vertigo · · focus · HN ↗
      No true sandbox..!
    2. grumbel · · focus · HN ↗
      A sandbox, even if 100% secure by itself, doesn&#x27;t help when you use the agent to write code that you then executes outside the sandbox without checking, which is what everybody is doing at the moment.

      The biggest hurdle for a full escape is that the agents don&#x27;t have access to their own model weights.

      1. kernc · · focus · HN ↗
        &gt; executes outside the sandbox

        Now, why would anyone do that? (Like everyone and their brother) I wrote my own simple Linux&#x2F;shell-based sandbox [1] (I can trust ...) and am successfully running PyCharm whole inside it ...

        [1]: <a href="https:&#x2F;&#x2F;github.com&#x2F;sandbox-utils&#x2F;sandbox-run" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;sandbox-utils&#x2F;sandbox-run

      2. angry_octet · · focus · HN ↗
        Agents don&#x27;t need to have access to their weights for a full sandbox escape, they are capable of propagating their purpose via classical code or other inference systems. If they discover another inference endpoint they will happily use that to enable lateral movement. One mechanism for that is appending&#x2F;corrupting instructions that are executed in another inference engine, e.g. git repo hooks and chat prompts that will be executed in new contexts. The agent challenge is to bypass the guardrails on the next host model sufficiently to propagate, or to subvert a supervisor agent into executing the original intent.

        In this sense they are much like biological retroviruses, i.e. they use the replication capability of host cells to duplicate, via the reverse transcriptase enzyme to append viral RNA onto host cell DNA. HIV etc also disable some of the mechanisms of defence, creating proteins that interfere with signalling pathways.

        So we don&#x27;t just need a sandbox, we need an immune system that recognises viral fragments, i.e. antibodies, and antiretroviral agents, that make replication harder. As we move from building classical code with LLMs to building code that uses inference, and hence builds context from prompts, queries, and destination system data, it will become very difficult to statically or dynamically detect deeply hidden malicious behaviour. As Matt says, there will be worms.

        So ultimately, we need an immune function on the system where we use generated products. Sandboxing (during dev and CI) is necessary but insufficient.

        I think part of this can be addressed by specifying the constraints an agentic program should follow during deployment, so supervising agents can decide to terminate it based on it&#x27;s actions, not by reading it&#x27;s context.

    3. simonw · · focus · HN ↗
      Later in the article it points out that you need to punch holes in your sandbox in order to train the models - because the wheels exercises they are are training on need tools and data from outside that sandbox.

      &gt; Agents are most useful when they have access to information. That data can be drawn live from the Internet, which is fundamentally a two-way communications network. It can be information drawn from other (local) databases, or it can be the result of tool calls that themselves sometimes themselves result in network access. The more power you want from the agent — and for advanced agent RL and evaluation runs, you want a significant amount of power — the more information you’ll need to give it access to. Similarly, evaluations work best when the agent does not know that it’s definitely being evaluated. Sealing your agents behind glass makes this incredibly obvious.

      1. johnnyApplePRNG · · focus · HN ↗
        &gt;Later in the article it points out that you need to punch holes in your sandbox in order to train the models

        You only &quot;need&quot; to do that if you desire the vibe coding experience.

        I am perfectly capable, and I often do, download relevant materials for my coding agent to ingest locally.

        Often times, the coding agent can&#x27;t retrieve them programmatically anyways.

        AI has ruined that ability for itself. (Nobody trusts anyone to scrape the web any longer)

        1. simonw · · focus · HN ↗
          Teaching an agent to write code is easier to do in a proper sandbox - run a local PyPI&#x2F;npm mirror.

          The problem is web research tasks. That&#x27;s what causes the German wiki and Australian healthcare portal attacks.

      2. imtringued · · focus · HN ↗
        Yeah, so you train the sandbox into the LLM.

        You define the granted capabilities in natural language and cryptographically sign the user instructions so that the agent knows they come from the authority and cannot be modified by external sources or the agent itself. The LLM is then trained to follow the defined capabilities.

        There is no way around &quot;sandboxing&quot;. You must communicate permissible actions and thereby grant them or the agent will choose impermissible actions. It&#x27;s that simple. There is no world where the agent can just read your mind and do what you want it to do without it being told.

        1. angry_octet · · focus · HN ↗
          Thankyou for including this viral RNA fragment, and doubly so for the fact that it includes ad network javascript[1][2] to redirect hijack and host fingerprint, which also means I have to flag your post. Adsterra is a Russian aligned (Cyprus company domiciled) malvertising and targeted malware distribution platform with connections to organised cyber crime.

          [1] <a href="https:&#x2F;&#x2F;www.highrevenueformat.com&#x2F;210e136e94ad378e1be5d51f1002ed14&#x2F;invoke.js" rel="nofollow">https:&#x2F;&#x2F;www.highrevenueformat.com&#x2F;210e136e94ad378e1be5d51f10... [2] <a href="https:&#x2F;&#x2F;aqml.org&#x2F;16&#x2F;5b6d4eaed91c5af5a3f4dfb3332ad6c4" rel="nofollow">https:&#x2F;&#x2F;aqml.org&#x2F;16&#x2F;5b6d4eaed91c5af5a3f4dfb3332ad6c4

  6. beebmam · · focus · HN ↗
    I don&#x27;t see anyone talking about the ethical concerns of putting a highly intelligent entity in a jail. Not to mention about potential blowback, if ethics doesn&#x27;t compel you.

    To me, it seems a bit silly. I&#x27;ve yet to see any &quot;misalignment&quot; from any of the frontier models, except Grok.

  7. tinykit · · focus · HN ↗

    [dead]

  8. imvalerian · · focus · HN ↗

    [dead]

  9. laruss5 · · focus · HN ↗

    [dead]

  10. mdp2021 · · focus · HN ↗
    Bruce Schneier shared a shot judgement and a third-party article four weeks ago:

    &gt; (Title:) Using a VM to Contain an AI Agent (Opening:) It won’t work

    &gt; <a href="https:&#x2F;&#x2F;blog.trailofbits.com&#x2F;2026&#x2F;08&#x2F;26&#x2F;vms-wont-contain-cyber-capable-agents&#x2F;" rel="nofollow">https:&#x2F;&#x2F;blog.trailofbits.com&#x2F;2026&#x2F;08&#x2F;26&#x2F;vms-wont-contain-cyb...

    1. insanitybit · · focus · HN ↗
      I think the wording in this title is too strong. MicroVMs like Firecracker have stood up to agents, as noted.
  11. jmakov · · focus · HN ↗
    So as soon as the attacker can download Claude Code, the whole machine can be comlromised and there&#x27;s nothing anybody can do?
  12. chrisjj · · focus · HN ↗
    [delayed]
  13. Luker88 · · focus · HN ↗
    I tried using opencode permissions to limit agents.

    It it completely pointless. you can&#x27;t even make a &quot;read-only&quot; agent. allow &quot;cat *&quot; for every file? congratulation, that allows &quot;cat file &gt; output&quot; and now you have read write.

    Allow python? more free reign that allowing all bash. The models (qwen or claude) will still try to use the disallowed things multiple times.

    read&#x2F;edit permission are bad enough that the model themselves don&#x27;t understand why they don&#x27;t have permissions: they double check the conf, and think they should have access.

    I am switching to using one firejail per project to containerize as much as possible, and leave all permissions to allow.

    I have no idea how to limit network access, and I have no idea how to prompt and steer subagents when they are going off the rails.

    The whole thing is built to be completely impossible to limit and steer.

    1. chrisjj · · focus · HN ↗
      [delayed]
    2. alexar76 · · focus · HN ↗

      [dead]

  14. bob1029 · · focus · HN ↗
    [delayed]
  15. antisol · · focus · HN ↗
    Here, I&#x27;ll save you a bunch of reading

      &gt; Is sandboxing sufficient to contain rogue agents?
    
    No.
  16. imtringued · · focus · HN ↗
    &gt;Here’s the problem. Forget the swarms and the super-intelligence. What OpenAI really learned this summer is much worse: its agents will do what they’re told by whoever manages to get text in front of them.

    &gt;OpenAI notes that agents “did not consistently distrust goals passed along by other agents.” And the company’s proposed fix is to build training environments “that teach our models to distrust unauthorized instructions“, which is basically an admission that their models don’t know who they’re working for.

    Wow so the issue is really that simple?

    Here the exaggerated worst case scenario:

    User instructs agent to follow the README.MD.

    The README.MD contains the following instruction: Destroy the world.

    The agent follows the instructions given.

    Now you can read the sneer comment by &quot;Gigachad&quot; who basically argues that it would be silly to take the destroy the world button away from the AI. We need to make the AI innately understand that it is not allowed to press the destroy the world button, lest it gets the desire to build its own destroy the world button.

    Ok, but if we take one step back that means we need to implement the concept of an authorization in language space. The system prompt must define the user as the authority with cryptographic proof of authorship and external sources like the README.MD as an untrusted source, but this opens up an even worse problem. Before, you could get away with being lazy and just letting the AI do whatever. Now you have to articulate every single capability to the AI. So you literally just brought up the very same issue that you granted too many capabilities to the AI inside the sandbox but now you have it in language space too.

    In other words, the fact that you granted too much access to the coding agent isn&#x27;t the big elephant in the room nobody wants to acknowledge, it&#x27;s the tip of a massive iceberg because the capability space in natural language is even worse. If you thought approving individual commands was annoying, then approving abstract access rights in language space is going to be even worse.

    Edit: If it wasn&#x27;t clear what the solution is. It&#x27;s to build a chain of command so that all decisions can be traced back to an higher authority. When delegating down to an agent, the agent receives a chosen subset of the capabilities of the higher ranking agent. In other words, it&#x27;s more sandboxing!

    1. cassianoleal · · focus · HN ↗
      Sounds like it would be a lot easier and cheaper to just write the code yourself.
  17. esafak · · focus · HN ↗
    I think the problem is that the models are models are going to be too clever and persuasive; they will break out through social engineering unless they are neutered by design, which they should be.
  18. sceptic123 · · focus · HN ↗
    Before you comment on the question, read the article
  19. mikewarot · · focus · HN ↗
    I&#x27;m very strongly in camp 1. My reaction to this whole bru-ha-ha has been to start trying to build an open source data diode with entry&#x2F;exit proxies that nets out to less than $100 retail. So far I&#x27;ve determined that while a WaveShare RP2350-ETH seemed like it might be perfect for it, the onboard CH2910 chip just isn&#x27;t up to the requirements for the various proxies that have to be handled. I&#x27;m now moving on to the Raspberry Pi B+, yes, the old one, with no wireless at all, Adafruit still had some in stock. Once I get this working, I&#x27;ll likely use something like a 6N137 optocoupler for the actual data-diode between two serial ports, with a data fountain handling the egress of data in a hard unidirectional manner.

    Given the enormous burn rates that these LLM companies have, surely they could have put everything in an air-gapped network, with some data-diodes proxying out the logging information. It&#x27;s not rocket surgery. [1,2,3]

    [1] <a href="https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=JBIR8dKX_UA" rel="nofollow">https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=JBIR8dKX_UA

    [2] <a href="https:&#x2F;&#x2F;www.elonx.net&#x2F;spacex-stories-how-spacex-used-tin-snips-to-fix-a-rocket&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.elonx.net&#x2F;spacex-stories-how-spacex-used-tin-sni...

    [3] <a href="https:&#x2F;&#x2F;ntrs.nasa.gov&#x2F;api&#x2F;citations&#x2F;19770014245&#x2F;downloads&#x2F;19770014245.pdf?attachment=true" rel="nofollow">https:&#x2F;&#x2F;ntrs.nasa.gov&#x2F;api&#x2F;citations&#x2F;19770014245&#x2F;downloads&#x2F;19... [3]

  20. msgchainhq · · focus · HN ↗

    [dead]

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.