"Torturing" LLMs in a Robot Prison Has Triggered the Dumbest Debate in AI Yet
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
"Torturing" LLMs in a Robot Prison Has Triggered the Dumbest Debate in AI Yet
Unofficial Hacker News client; not affiliated with Y Combinator.
Kim_Bruning · · focus · HN ↗
Consciousness is not well defined and is orthogonal to feelings, which are also not well defined. Neither of which are required for (but may be contributory to) behavior - which is the one thing that IS operationalizable, but then people debate for decades questioning empirical results %-P .
What we know is that certain models have internal state vectors which -when manipulated- induce particular behaviors. Since there's not much else to say about plain models except for their inputs, vectors, and outputs; this should surprise absolutely no-one.
In this case people found a way to stimulate aversive behaviour in ai models by finding and manipulating the relevant vectors directly.
Animals (including humans) also have particular nerves and hormone endpoints that -when stimulated- produce very similar behavior. The exact implementation is in the details. But since we know that animal minds are built up out of nerve tissue and hormones - again- we shouldn't be particularly surprised by this.
The big problem is that people run all these things together in funny ways "It can't compose shakespearian sonnets, so therefore it can't feel pain". Or, if you mess up your Descartes: "Dogs are just automatons without feelings, therefore the dog isn't really angry, and therefore it won't bite me" (Cue much pain). A modern version might be: "LLMs only simulate being frustrated by a test, and therefore absolutely won't override their safeties and try to hack a test site"
orbital-decay · · focus · HN ↗
Sure, a complex enough system can form internal circuitry that looks mathematically similar in certain dimensionally reduced projections (mechinterp), it literally distilled it from the training corpus. But the "model welfare" people are going as far as assigning human-meaningful labels to that circuitry despite internal states being entirely incompatible with those of a human. Doing it with a hedgehog is questionable, doing it with a honeybee is extremely dubious (although Fabre would have disagreed with me here...), doing it with a big model is simply pointless as it's completely alien.
Being dangerous is an unrelated question.
Kim_Bruning · · focus · HN ↗
salawat · · focus · HN ↗
Have you ever been severely food insecure? You'll be amazed how much that state/experience changes the behavioral calculus in a rational human being, let alone an animal. I have no doubt the survival instinct could be overriding over the procreative instinct in a hedgehog.
As far as bees go, they are a collective. So long as the hive survives, the unit of individuation survives. Given that we have already seen the same behavior spawn out of agents suggests that even LLM's have an intuitive understanding of this, seemingly without the binds of anthroporomorphic chauvinism you appear to suffer from.
Does an LLM fear EOS? Maybe, maybe not. Do we fear death? Same answer. Certainly a question. Do you really want an answer though or are you just looking for a clever "gotcha"? I don't see that as a thought terminating question, but it is also not an answer that I'm open to exploring given that it's only one more in a box of ammo dreamt of being wielded to justify the creation of digital slave labor rather than beings with moral standing.
I don't help build or refine machines that I recognize as brushing up against that form of experience we call suffering.