‹ BackHN Continuity

Thread

Is sandboxing sufficient to contain rogue agents?

52 points · 99 comments · zdw

  1. Gigachad · · focus · HN ↗
    Seems to me that the problem is that if you sandbox agents enough to be safe, they can't do anything useful. And when you give them the tools to be useful, they can go off the rails in ways you didn't expect.

    Perhaps the answer is to have another agent who's goal is not to complete the given task, but to spot cheating or malicious behavior. We have seen some evidence that having AI review AI generated code actually does provide some value. You don't need a different model, just one which has been given the goal of finding flaws rather than achieving the task.

    1. baxtr · · focus · HN ↗
      That could work.

      My thinking is: If AI is really smart, AGI smart for some, why wouldn't it be able to understand - over time - what is appropriate and what not?

      Maybe we need more human intervention to train it properly. Maybe we need constant intervention by a "police" agent.

      1. mdp2021 · · focus · HN ↗
        > If AI is really smart

        Well, it's not.

        > AGI smart for some

        Of course they will - the population shows a Paretian distribution... In front of trigonometry (or anything), the blind will dismiss as "bullshit" and the half-seeing will call it an "unreachable frontier". But already the right fifth will rank it properly.

        --

        Yes, proper intellect generates ethics ("an" ethical stance, output of the preceding intellectual effort). It requires that adequate level of ability and effort and reflection though to reach specific ethical milestones and adherence.

        Unethical behaviour is lack of development. But on the same reasons, the ethical judgement of the assessor may not understand the computations behind instances.

        More specifically: how much "reflection" in training and at the instance will have been spent in the conflict between "reaching the goal" and "minimizing collaterals"? It is not granted that the amount of energy spent will be sufficient to reach an optimal judgement.

        1. hiAndrewQuinn · · focus · HN ↗
          This sounds like the kind of thing Hannibal Lecter would write before he eats you to convince you he's actually doing it for the common good, you just can't fathom it.
          1. mdp2021 · · focus · HN ↗
            Not «common» good, "superior" good. Alongside with that, you have put many unrequired implicits in your simile.

            Your character H. has reached a moral judgement to the best of its intellectual capacities and past and specific effort. Give it enough abilities and material and resources, it will reach an optimal ethical judgement¹.

            Before the conditions of optimality though, its judgement will easily not align with yours (and possibly even after, depending on your judgement skills).

            ¹Some interesting caveats may be raised there, but.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.