"Torturing" LLMs in a Robot Prison Has Triggered the Dumbest Debate in AI Yet
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
"Torturing" LLMs in a Robot Prison Has Triggered the Dumbest Debate in AI Yet
Unofficial Hacker News client; not affiliated with Y Combinator.
Kim_Bruning · · focus · HN ↗
Consciousness is not well defined and is orthogonal to feelings, which are also not well defined. Neither of which are required for (but may be contributory to) behavior - which is the one thing that IS operationalizable, but then people debate for decades questioning empirical results %-P .
What we know is that certain models have internal state vectors which -when manipulated- induce particular behaviors. Since there's not much else to say about plain models except for their inputs, vectors, and outputs; this should surprise absolutely no-one.
In this case people found a way to stimulate aversive behaviour in ai models by finding and manipulating the relevant vectors directly.
Animals (including humans) also have particular nerves and hormone endpoints that -when stimulated- produce very similar behavior. The exact implementation is in the details. But since we know that animal minds are built up out of nerve tissue and hormones - again- we shouldn't be particularly surprised by this.
The big problem is that people run all these things together in funny ways "It can't compose shakespearian sonnets, so therefore it can't feel pain". Or, if you mess up your Descartes: "Dogs are just automatons without feelings, therefore the dog isn't really angry, and therefore it won't bite me" (Cue much pain). A modern version might be: "LLMs only simulate being frustrated by a test, and therefore absolutely won't override their safeties and try to hack a test site"
onlyrealcuzzo · · focus · HN ↗
This is orthogonal to whether they "feel" "pain".
They could be dangerous or not dangerous whether or not they feel pain.
A chess engine doesn't need to "feel angry" to annihilate me - I'm terrible at Chess.
An automated missile system doesn't need to be smart to wipe out humanity, just misaligned goals.
It seems like you made a good argument, and then lumped on a conclusion that defeats it...
Kim_Bruning · · focus · HN ↗
That is rather my point.
Your chess engine doesn't need to "feel pain", but most of them do apply some form of weighted tree search to find the next most optimal move, right?
It's sort of a similar thing: LLMs do something a bit more high dimensional, and have a lot more weighting vectors while computing the next most optimal token (fsvo optimal).
For instance, the experiment at hand demonstrates the existence of 'pain vectors'. They do so by altering them and observing whether there is an effect. That's pretty scientific.
There's also a 'desperation vector' that was studied by Anthropic interpretability folks earlier; that one is pretty much predictive of cheating.
I'm looking forward to seeing interpretability papers on other such vectors too.
orbital-decay · · focus · HN ↗
Kim_Bruning · · focus · HN ↗
Whether that's 'real suffering' or merely a convincing simulation is a job for the philosophers.
(I do have my own opinion, mind, and it's not what you might expect O:-) But the Overton window isn't there. A lot of people don't realize these vectors exist at all yet.)
Edit: On rereading, it might seem like I'm dodging the question. I'm really just trying to stick to my core points: A) the vectors exist B) they have a causal role in behavior, irrespective of the moral patienthood question.