"Torturing" LLMs in a Robot Prison Has Triggered the Dumbest Debate in AI Yet
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
"Torturing" LLMs in a Robot Prison Has Triggered the Dumbest Debate in AI Yet
Unofficial Hacker News client; not affiliated with Y Combinator.
Svip · · focus · HN ↗
As far as I am aware, the GPUs don't work harder, don't heat up more, and nothing in their mathematical equations are different, when the result of these vector calculations produce text, that to humans look like someone in pain. Whereas real pain is objectively observable, LLMs' pain - so far at least - is entirely subjectively observable.
That being said, I'm a fan of Star Trek's depiction of Data in TNG and the Doctor in VOY fighting for their rights to be treated as equals to the biological crew; and while I support their plight, as far as I can remember, the shows never made a claim that their pain was observable by instruments, thus making the judgement an entirely subjective decision (particularly in the case of the Doctor). Though I'd be glad to be corrected in this matter.
Kim_Bruning · · focus · HN ↗
The experiment under discussion demonstrates exactly that: changing the vectors induces outputs corresponding to what we would see as utterances of pain.
I've noodled with some of this myself to the point that I'm mostly convinced; but there's at least 2 papers I'm aware of on this:
* <a href="https://arxiv.org/abs/2604.07729" rel="nofollow">https://arxiv.org/abs/2604.07729 Nicholas Sofroniew et al "Emotion Concepts and their Function in a Large Language Model"
* <a href="https://arxiv.org/abs/2609.16247" rel="nofollow">https://arxiv.org/abs/2609.16247 Valen Tagliabue et al "The Pain Axis: LLMs Represent Self-Directed Harm and Act on It"