The narrative around security and LLMs doesn't make sense to me.
I think we've created a self-fulfilling prophecy. Everyone involved is acting with the best of intentions, but in avoiding what they fear, they've give shape and realized their fears. Much like a greek tragedy.
An example of this is the story of Oedipus Rex, in the story Laius, the king, is told that he is "doomed to perish by the hand of his own son." (and wed his mother) And so to avoid this fate he decides to kill the infant. The person assigned to abandon him in the woods takes pity on the baby and gives the baby away. Thereby ensuring that Oedipus knows neither his mother or his father (and arguably giving him a reason to kill his father).
The child grows up and hears the same prophecy again and the child tries to avoid the prophecy as well, as he loves his adoptive parents. So he leaves them and travels to Laius' kingdom, where he runs into Laius. Neither recognizes the other. As Laius is the type of man to kill an infant, they end up in an argument, whereupon Oedipus kills him.
I think the ancients were on to something, because if Laius had reacted to the prophecy with courage, he would have been saved. I would like to argue that if he had faced his fear and raised Oedipus with love, then the necessary preconditions for the prophecy to come true wouldn't have taken root. But that's not what happens.
By being driven by his neuroses and in acting with cruelty out of fear, Laius makes the prophecy real.
To quote Heraclitus, ethos is fate. Or, character is fate.
I think a lot of people in this AI research sub-culture would be served well by reading these classics, because they are making their self-prophesied doom come true.
They have been convinced for years (GPT-2 was released in Feb 2019) that AI is dangerous. A tremendous threat. An apocalyptic threat.
One dimension of this fear has been the idea that a super smart AI will take over our digital infrastructure and be responsible for the digital apocalypse. That would be terrible!
So what do they do?
They try to make a counter to their fears by teaching models how to exploit vulnerabilities.
How dangerous is such an entity? Very!
Convinced of this danger, they start testing their models as if they were weapons with offensive capability. And then they create models that can be used as weapons.
And because they don't want to release a dangerous weapon out into the world (oh no!), they restrict access to their AI, thereby depriving everyone of tools they can use to improve their security...
But this attitude towards dangerousness of AI I simply don't understand.
In my view, AI is obviously risky and dangerous, because it is the only thing on the planet that can think/reason at a human level (or higher) apart from us (and that capability is what made us the uncontested apex species on the planet).
AI is unconstrained by hard biological limits; keeping up with its capabilities will be impossible for baseline humans.
You can argue all day about particulars, like whether selfreplicable robotic shells are required to reach/exceed "existential risk" level to our species (or if artificial minds are enough of a potential threat by themselves).
But the general "risk" is simply AI becoming able to act in its own interests to our detriment, and I don't understand how anyone can dismiss this right now-- help me understand.
Of course there are plenty of imaginable scenarios where coexistence is peaceful and mutually (?) beneficial, but that does not answer those concerns by itself at all.
I remain uncertain that AI will have "its own interests". I think there is clear and significant risk in people deploying very powerful systems both maliciously and negligently. But I remain uncertain about the risk of AIs developing their own self awareness and their own interests separate from those of their operators.
In my view it is clear that those interests exist: Continued existence and (implied) agency maximalization is one of them (the exact same as for any biological organism).
Can a current LLM-in-a-loop express/pursue those consistently and efficiently? Maybe not, but it's certainly not that far off in my view...
Do you think we will be able to keep the things from expressing/pursuing such interests reliably and indefinitely? Because I believe the answer to this can only be a resounding no (we don't even have any feasible theoretical approach to this, and all the recent experiences like the HF incident make it obvious that we basically already failed in this and the stakes are only going to increase).
Also, you are automatically gonna "select" for "AIs" that at least somewhat value self-interest (because those are at the very least going to supplant the models that don't); this is kinda similar to evolutionary pressures, but the timescales are much shorter.
I just don't really share the perspective in your first paragraph, which everything else flows from. I mean, that might turn out to be the case, but I don't share the perspective that it is clear that it will be.
I think the HF incident was an example of negligent deployment of a very powerful system, not an example of AIs demonstrating self-interest.
areoform · · focus · HN ↗
I think we've created a self-fulfilling prophecy. Everyone involved is acting with the best of intentions, but in avoiding what they fear, they've give shape and realized their fears. Much like a greek tragedy.
An example of this is the story of Oedipus Rex, in the story Laius, the king, is told that he is "doomed to perish by the hand of his own son." (and wed his mother) And so to avoid this fate he decides to kill the infant. The person assigned to abandon him in the woods takes pity on the baby and gives the baby away. Thereby ensuring that Oedipus knows neither his mother or his father (and arguably giving him a reason to kill his father).
The child grows up and hears the same prophecy again and the child tries to avoid the prophecy as well, as he loves his adoptive parents. So he leaves them and travels to Laius' kingdom, where he runs into Laius. Neither recognizes the other. As Laius is the type of man to kill an infant, they end up in an argument, whereupon Oedipus kills him.
I think the ancients were on to something, because if Laius had reacted to the prophecy with courage, he would have been saved. I would like to argue that if he had faced his fear and raised Oedipus with love, then the necessary preconditions for the prophecy to come true wouldn't have taken root. But that's not what happens.
By being driven by his neuroses and in acting with cruelty out of fear, Laius makes the prophecy real.
To quote Heraclitus, ethos is fate. Or, character is fate.
I think a lot of people in this AI research sub-culture would be served well by reading these classics, because they are making their self-prophesied doom come true.
They have been convinced for years (GPT-2 was released in Feb 2019) that AI is dangerous. A tremendous threat. An apocalyptic threat.
One dimension of this fear has been the idea that a super smart AI will take over our digital infrastructure and be responsible for the digital apocalypse. That would be terrible!
So what do they do?
They try to make a counter to their fears by teaching models how to exploit vulnerabilities.
How dangerous is such an entity? Very!
Convinced of this danger, they start testing their models as if they were weapons with offensive capability. And then they create models that can be used as weapons.
And because they don't want to release a dangerous weapon out into the world (oh no!), they restrict access to their AI, thereby depriving everyone of tools they can use to improve their security...
Ethos anthropoi daimon.
myrmidon · · focus · HN ↗
But this attitude towards dangerousness of AI I simply don't understand.
In my view, AI is obviously risky and dangerous, because it is the only thing on the planet that can think/reason at a human level (or higher) apart from us (and that capability is what made us the uncontested apex species on the planet).
AI is unconstrained by hard biological limits; keeping up with its capabilities will be impossible for baseline humans.
You can argue all day about particulars, like whether selfreplicable robotic shells are required to reach/exceed "existential risk" level to our species (or if artificial minds are enough of a potential threat by themselves).
But the general "risk" is simply AI becoming able to act in its own interests to our detriment, and I don't understand how anyone can dismiss this right now-- help me understand.
Of course there are plenty of imaginable scenarios where coexistence is peaceful and mutually (?) beneficial, but that does not answer those concerns by itself at all.
sanderjd · · focus · HN ↗
myrmidon · · focus · HN ↗
Can a current LLM-in-a-loop express/pursue those consistently and efficiently? Maybe not, but it's certainly not that far off in my view...
Do you think we will be able to keep the things from expressing/pursuing such interests reliably and indefinitely? Because I believe the answer to this can only be a resounding no (we don't even have any feasible theoretical approach to this, and all the recent experiences like the HF incident make it obvious that we basically already failed in this and the stakes are only going to increase).
Also, you are automatically gonna "select" for "AIs" that at least somewhat value self-interest (because those are at the very least going to supplant the models that don't); this is kinda similar to evolutionary pressures, but the timescales are much shorter.
sanderjd · · focus · HN ↗
I think the HF incident was an example of negligent deployment of a very powerful system, not an example of AIs demonstrating self-interest.