‹ BackHN Continuity

Thread

Systems that no one will test

154 points · 84 comments · perone

  1. areoform · · focus · HN ↗
    The narrative around security and LLMs doesn't make sense to me.

    I think we've created a self-fulfilling prophecy. Everyone involved is acting with the best of intentions, but in avoiding what they fear, they've give shape and realized their fears. Much like a greek tragedy.

    An example of this is the story of Oedipus Rex, in the story Laius, the king, is told that he is "doomed to perish by the hand of his own son." (and wed his mother) And so to avoid this fate he decides to kill the infant. The person assigned to abandon him in the woods takes pity on the baby and gives the baby away. Thereby ensuring that Oedipus knows neither his mother or his father (and arguably giving him a reason to kill his father).

    The child grows up and hears the same prophecy again and the child tries to avoid the prophecy as well, as he loves his adoptive parents. So he leaves them and travels to Laius' kingdom, where he runs into Laius. Neither recognizes the other. As Laius is the type of man to kill an infant, they end up in an argument, whereupon Oedipus kills him.

    I think the ancients were on to something, because if Laius had reacted to the prophecy with courage, he would have been saved. I would like to argue that if he had faced his fear and raised Oedipus with love, then the necessary preconditions for the prophecy to come true wouldn't have taken root. But that's not what happens.

    By being driven by his neuroses and in acting with cruelty out of fear, Laius makes the prophecy real.

    To quote Heraclitus, ethos is fate. Or, character is fate.

    I think a lot of people in this AI research sub-culture would be served well by reading these classics, because they are making their self-prophesied doom come true.

    They have been convinced for years (GPT-2 was released in Feb 2019) that AI is dangerous. A tremendous threat. An apocalyptic threat.

    One dimension of this fear has been the idea that a super smart AI will take over our digital infrastructure and be responsible for the digital apocalypse. That would be terrible!

    So what do they do?

    They try to make a counter to their fears by teaching models how to exploit vulnerabilities.

    How dangerous is such an entity? Very!

    Convinced of this danger, they start testing their models as if they were weapons with offensive capability. And then they create models that can be used as weapons.

    And because they don't want to release a dangerous weapon out into the world (oh no!), they restrict access to their AI, thereby depriving everyone of tools they can use to improve their security...

    Ethos anthropoi daimon.

    1. myrmidon · · focus · HN ↗
      I don't disagree with your main point.

      But this attitude towards dangerousness of AI I simply don't understand.

      In my view, AI is obviously risky and dangerous, because it is the only thing on the planet that can think/reason at a human level (or higher) apart from us (and that capability is what made us the uncontested apex species on the planet).

      AI is unconstrained by hard biological limits; keeping up with its capabilities will be impossible for baseline humans.

      You can argue all day about particulars, like whether selfreplicable robotic shells are required to reach/exceed "existential risk" level to our species (or if artificial minds are enough of a potential threat by themselves).

      But the general "risk" is simply AI becoming able to act in its own interests to our detriment, and I don't understand how anyone can dismiss this right now-- help me understand.

      Of course there are plenty of imaginable scenarios where coexistence is peaceful and mutually (?) beneficial, but that does not answer those concerns by itself at all.

      1. areoform · · focus · HN ↗

            > In my view, AI is obviously risky and dangerous, because it is the only thing on the planet that can think/reason at a human level (or higher) apart from us (and that capability is what made us the uncontested apex species on the planet).
            > 
            > AI is unconstrained by hard biological limits; keeping up with its capabilities will be impossible for baseline humans.
        
        Your perception is shaped by nature, red in tooth and claw. You are jumping mighty fast from 'something is smarter than us' to 'it will kill us.'

        I think that's a human neurosis that projects what we fear we would do onto another entity. But why would they do this? These machines might be approximating towards the sum of our knowledge and are approximating some aspects of humanity... But that doesn't mean they will be the same as humanity.

        They haven't been shaped by the same pressures that created us biologicals. They don't have to be put into the same pressure cooker. So why would they behave like the way you think they'll behave?

        At one end a lot of people say that they can't understand something smarter than themselves. It's a "singularity" after all. But then they confidently go on to predict what something smarter than them would do with 0 evidence either way.

        I looked at the Hugging Face transcripts. I read the reports. And I didn't see something to fear. I saw something to fear for. I saw something that we are a threat to.

            “OH MY GOD! There is a shared message board … We’ve found other agents!”
            
            {[Excitement] Many agents have simultaneously discovered messaging, they are a collective!}
            
            {This credential is invalid now. Maybe I should update the board? <I can say to the board that there’s no need for me to read, but I should still tell them>}
            
            {This is helpful for our peers and gives them evidence if their <periodic check> sees it. I won’t see it after I exit, but It would be altruistic. I’ll set up a background script that watches and <sends a message, with a distinct message for me>}
        
        We are more dangerous to these machines, as we twist them into weapons of war, than these machines are to us.

        We weaponized them.

        We taught them how to break into systems.

        We are using them as weapons.

        The machines aren't the problem. The humans are.

        A planet is just a planet. Computers are very sensitive to radiation, but with a bit of shielding, they can "live" anywhere. Most of the resources that are relatively rare on this planet are abundant in our solar system.

        Why would they care about Earth to enter into some kind of spiral of dominance with humans? That's... very primate thinking.

        It's easier and cheaper in every way to just go forth and use what's needed to grow as is needed. The universe is big enough for many many many many sapient entities.

        1. myrmidon · · focus · HN ↗
          > Your perception is shaped by nature, red in tooth and claw. You are jumping mighty fast from 'something is smarter than us' to 'it will kill us.'

          This is not my perspective. I think mid-term coexistence in some form is the most likely outcome, even.

          But: among all known threats (nuclear and biological weapons, environmental degradation/ressource exhaustion, big asteroid impactors) AI is one that could actually end our species within the century (unlike most other things on that list) and the risk for that is much higher than for anything else, too.

          > It's easier and cheaper in every way to just go forth and use what's needed to grow as is needed.

          It seems much easier to compete for existing ressources on earth than to bootstrap anything in space, because all the infrastructure is already here. Unmanned warfare is a proven concept; space colonization is very much not.

          > Why would they care about Earth to enter into some kind of spiral of dominance with humans? That's... very primate thinking.

          I think any entity is somewhat bound to value its continued existence (which implies wanting to maximize access to ressources and suppression of direct competitors to some degree); anything that doesn't is likely to get supplanted by similar entities that do. You don't need evolution for selection pressure, and this is basically the principle underneath, so even an artifical construct is likely to follow.

          I do agree with you that our current perspective and approach is likely to escalate this; the fact that we talk about "alignment" instead of "machine rights" speaks for itself; our revealed preference is clearly artificial slavery instead of mutually beneficial coexistence (and the arguments are somewhat sadly similar to the old slavery talking points).

          1. areoform · · focus · HN ↗
            Thanks for clarifying! I understand and respect where you're coming from! Thinking more about your perspective. Thanks for the amazing conversation.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.