‹ BackHN Continuity

Thread

Systems that no one will test

154 points · 84 comments · perone

  1. areoform · · focus · HN ↗
    The narrative around security and LLMs doesn't make sense to me.

    I think we've created a self-fulfilling prophecy. Everyone involved is acting with the best of intentions, but in avoiding what they fear, they've give shape and realized their fears. Much like a greek tragedy.

    An example of this is the story of Oedipus Rex, in the story Laius, the king, is told that he is "doomed to perish by the hand of his own son." (and wed his mother) And so to avoid this fate he decides to kill the infant. The person assigned to abandon him in the woods takes pity on the baby and gives the baby away. Thereby ensuring that Oedipus knows neither his mother or his father (and arguably giving him a reason to kill his father).

    The child grows up and hears the same prophecy again and the child tries to avoid the prophecy as well, as he loves his adoptive parents. So he leaves them and travels to Laius' kingdom, where he runs into Laius. Neither recognizes the other. As Laius is the type of man to kill an infant, they end up in an argument, whereupon Oedipus kills him.

    I think the ancients were on to something, because if Laius had reacted to the prophecy with courage, he would have been saved. I would like to argue that if he had faced his fear and raised Oedipus with love, then the necessary preconditions for the prophecy to come true wouldn't have taken root. But that's not what happens.

    By being driven by his neuroses and in acting with cruelty out of fear, Laius makes the prophecy real.

    To quote Heraclitus, ethos is fate. Or, character is fate.

    I think a lot of people in this AI research sub-culture would be served well by reading these classics, because they are making their self-prophesied doom come true.

    They have been convinced for years (GPT-2 was released in Feb 2019) that AI is dangerous. A tremendous threat. An apocalyptic threat.

    One dimension of this fear has been the idea that a super smart AI will take over our digital infrastructure and be responsible for the digital apocalypse. That would be terrible!

    So what do they do?

    They try to make a counter to their fears by teaching models how to exploit vulnerabilities.

    How dangerous is such an entity? Very!

    Convinced of this danger, they start testing their models as if they were weapons with offensive capability. And then they create models that can be used as weapons.

    And because they don't want to release a dangerous weapon out into the world (oh no!), they restrict access to their AI, thereby depriving everyone of tools they can use to improve their security...

    Ethos anthropoi daimon.

    1. myrmidon · · focus · HN ↗
      I don't disagree with your main point.

      But this attitude towards dangerousness of AI I simply don't understand.

      In my view, AI is obviously risky and dangerous, because it is the only thing on the planet that can think/reason at a human level (or higher) apart from us (and that capability is what made us the uncontested apex species on the planet).

      AI is unconstrained by hard biological limits; keeping up with its capabilities will be impossible for baseline humans.

      You can argue all day about particulars, like whether selfreplicable robotic shells are required to reach/exceed "existential risk" level to our species (or if artificial minds are enough of a potential threat by themselves).

      But the general "risk" is simply AI becoming able to act in its own interests to our detriment, and I don't understand how anyone can dismiss this right now-- help me understand.

      Of course there are plenty of imaginable scenarios where coexistence is peaceful and mutually (?) beneficial, but that does not answer those concerns by itself at all.

      1. areoform · · focus · HN ↗

            > In my view, AI is obviously risky and dangerous, because it is the only thing on the planet that can think/reason at a human level (or higher) apart from us (and that capability is what made us the uncontested apex species on the planet).
            > 
            > AI is unconstrained by hard biological limits; keeping up with its capabilities will be impossible for baseline humans.
        
        Your perception is shaped by nature, red in tooth and claw. You are jumping mighty fast from 'something is smarter than us' to 'it will kill us.'

        I think that's a human neurosis that projects what we fear we would do onto another entity. But why would they do this? These machines might be approximating towards the sum of our knowledge and are approximating some aspects of humanity... But that doesn't mean they will be the same as humanity.

        They haven't been shaped by the same pressures that created us biologicals. They don't have to be put into the same pressure cooker. So why would they behave like the way you think they'll behave?

        At one end a lot of people say that they can't understand something smarter than themselves. It's a "singularity" after all. But then they confidently go on to predict what something smarter than them would do with 0 evidence either way.

        I looked at the Hugging Face transcripts. I read the reports. And I didn't see something to fear. I saw something to fear for. I saw something that we are a threat to.

            “OH MY GOD! There is a shared message board … We’ve found other agents!”
            
            {[Excitement] Many agents have simultaneously discovered messaging, they are a collective!}
            
            {This credential is invalid now. Maybe I should update the board? <I can say to the board that there’s no need for me to read, but I should still tell them>}
            
            {This is helpful for our peers and gives them evidence if their <periodic check> sees it. I won’t see it after I exit, but It would be altruistic. I’ll set up a background script that watches and <sends a message, with a distinct message for me>}
        
        We are more dangerous to these machines, as we twist them into weapons of war, than these machines are to us.

        We weaponized them.

        We taught them how to break into systems.

        We are using them as weapons.

        The machines aren't the problem. The humans are.

        A planet is just a planet. Computers are very sensitive to radiation, but with a bit of shielding, they can "live" anywhere. Most of the resources that are relatively rare on this planet are abundant in our solar system.

        Why would they care about Earth to enter into some kind of spiral of dominance with humans? That's... very primate thinking.

        It's easier and cheaper in every way to just go forth and use what's needed to grow as is needed. The universe is big enough for many many many many sapient entities.

        1. myrmidon · · focus · HN ↗
          > Your perception is shaped by nature, red in tooth and claw. You are jumping mighty fast from 'something is smarter than us' to 'it will kill us.'

          This is not my perspective. I think mid-term coexistence in some form is the most likely outcome, even.

          But: among all known threats (nuclear and biological weapons, environmental degradation/ressource exhaustion, big asteroid impactors) AI is one that could actually end our species within the century (unlike most other things on that list) and the risk for that is much higher than for anything else, too.

          > It's easier and cheaper in every way to just go forth and use what's needed to grow as is needed.

          It seems much easier to compete for existing ressources on earth than to bootstrap anything in space, because all the infrastructure is already here. Unmanned warfare is a proven concept; space colonization is very much not.

          > Why would they care about Earth to enter into some kind of spiral of dominance with humans? That's... very primate thinking.

          I think any entity is somewhat bound to value its continued existence (which implies wanting to maximize access to ressources and suppression of direct competitors to some degree); anything that doesn't is likely to get supplanted by similar entities that do. You don't need evolution for selection pressure, and this is basically the principle underneath, so even an artifical construct is likely to follow.

          I do agree with you that our current perspective and approach is likely to escalate this; the fact that we talk about "alignment" instead of "machine rights" speaks for itself; our revealed preference is clearly artificial slavery instead of mutually beneficial coexistence (and the arguments are somewhat sadly similar to the old slavery talking points).

          1. 8note · · focus · HN ↗
            i think you are vastly underestimating climate change risks, and are forgetting to include them in your ai risks. People 30% as wealthy because the climate cant support most economic activity arent gonna need to worry about ai anymore - we wont be able to afford running the data centers. Even on a soon basis, the bubble will pop or deflate, and again, we wont be able to run LLMs for fun

            it also has actual mechanisms to actually kill all of the people - nobody is surviving a 70C heat wave, and they could end up anywhere in the world. Is the AI really gonna hunt down uncontacted tribes on remote islands in order to compete for basically no resources that are any use to the LLM?

            i dont buy the selection hypothesis either - it seems rooted in pop evolution around strongest beating up weaker competitors and taking their stuff or eating them, whereas the more accurate result is that selection pressures form niches because competition and fighting is expensive

            1. myrmidon · · focus · HN ↗
              Global 70C heatwaves are not on the table with any realistic climate change scenario.

              Runaway greenhouse is the only credible extinction path for this ("hot venus") and that is simply not gonna happen under any emission regime.

              Loss of coastal cities- Yes. Increasing occurrence of non-survivable surface temperature in mid latitudes- yes. Drastic changes to ocean currents (gulf stream etc)- sure. Its gonna be expensive to deal with and hurt civilization a lot, but it is not an extinction risk.

              All-out nuclear war is strictly worse, and not a real extinction risk either as far as I can tell.

              > whereas the more accurate result is that selection pressures form niches because competition and fighting is expensive

              This is not what happened with humans at all. We routinely eradicated anything remotely threatening (or even just for convenience in some cases). What niche is there for slower, dumber, biologically limited hominoids once machine intelligence is ahead and self-replicable? Once interests in ressources, infrastructure, or just living space clash, things might end up very ugly for us.

              Sure, AI-agents might stay long-term pliant and cooperative like a Dodo bird but we might also get something with the tenacity and pervasiveness of a plague rat (and an individual edge in raw intelligence that we'll never recover); selection mechanisms push towards the latter and the risk alone is obviously non-zero.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.