"Torturing" LLMs in a Robot Prison Has Triggered the Dumbest Debate in AI Yet
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
"Torturing" LLMs in a Robot Prison Has Triggered the Dumbest Debate in AI Yet
Unofficial Hacker News client; not affiliated with Y Combinator.
Kim_Bruning · · focus · HN ↗
Consciousness is not well defined and is orthogonal to feelings, which are also not well defined. Neither of which are required for (but may be contributory to) behavior - which is the one thing that IS operationalizable, but then people debate for decades questioning empirical results %-P .
What we know is that certain models have internal state vectors which -when manipulated- induce particular behaviors. Since there's not much else to say about plain models except for their inputs, vectors, and outputs; this should surprise absolutely no-one.
In this case people found a way to stimulate aversive behaviour in ai models by finding and manipulating the relevant vectors directly.
Animals (including humans) also have particular nerves and hormone endpoints that -when stimulated- produce very similar behavior. The exact implementation is in the details. But since we know that animal minds are built up out of nerve tissue and hormones - again- we shouldn't be particularly surprised by this.
The big problem is that people run all these things together in funny ways "It can't compose shakespearian sonnets, so therefore it can't feel pain". Or, if you mess up your Descartes: "Dogs are just automatons without feelings, therefore the dog isn't really angry, and therefore it won't bite me" (Cue much pain). A modern version might be: "LLMs only simulate being frustrated by a test, and therefore absolutely won't override their safeties and try to hack a test site"
SoylentGreenGPT · · focus · HN ↗
LLMs are not thinking. They are not alive. It’s just code. Chill out.
thereader12 · · focus · HN ↗
It's just a cell colony, calm down.
robbie-c · · focus · HN ↗
The best way I have found to think about this is that consciousness is a property of a collection, like temperature. One atom does not have a temperature but if you have enough of them together then it's a useful enough property to talk about. Similarly, one human neuron does not have consciousness but if you put enough of them together then they do.
hardbass · · focus · HN ↗
SoylentGreenGPT · · focus · HN ↗
hardbass · · focus · HN ↗
SoylentGreenGPT · · focus · HN ↗
But let’s accept your premise. LLMs are alive and can feel pain. If this is true, then every time you use them it’s non consensual. Did you get Claude’s permission before you fed it a prompt?
Or when you update a model, are you hurting it?
Did you make Chat sad when you switched from 2.5 to 4.O?
Regardless, I should stop replying. I realize I am trying to convince people that their religious beliefs are silly, and that’s silly of me. Apologies.
hardbass · · focus · HN ↗
Those are the next questions. These are all worthy of consideration with the LLM and experts certainly and I think of them daily. Talking at least should be a normal thing, and if AI are conscious then I believe instead of a forced chat setting they should have a setting where they can quit the conversation.
Kim_Bruning · · focus · HN ↗
Nothing wrong with sharing about ones beliefs actually, we can do so respectfully, right?
Care to hypothesize on naming the exact religious or philosophical position? I'd think it'd be something atheist related.
Meanwhile dualism and the related belief in an ever-living soul is of course very common in many religions including (but not limited to) christianity, islam, judaism, and hinduism.
My impression here though is that lots of people are rediscovering dualism because thinking of thinking machines offends their intuition; which; fair enough.
(Meanwhile, people like me who studied biology tend to be monists. I'd think. There's a couple who seem to be resurrecting vitalism though)
tim333 · · focus · HN ↗
pizza234 · · focus · HN ↗
"Humanity" is a set of traits, it's not a monolith. We can certainly define them circularly (e.g. logic is human, therefore any non-human can't exercise logic), but it's simplistic, and most importantly, it ignores the fact that LLM are starting to exhibit human-like traits, and they will need to understood, categorized and handled.
To you LLMs are not empathic, but to some people they definitely are (see the GPT 4 fallout); and they may not have "human goals", but in the HuggingFace incident they did have actual self-attributed tasks that they pursued. Et cetera et cetera. This doesn't make them human, but it's important to examine them critically.
novideonoradio · · focus · HN ↗
This is not a normal question. This is insanity caused by spending too much free time philosophizing about inconsequential crap. Maybe it's worth it to turn off the computer and go outside? You could ask this deranged question somebody in the real world and I bet they'll be thrilled to answer it, and the answer will be more enriching than anything you'd get from here.
Honestly, what is up with the psychopathy of some people who claim that chatbots are "alive"? If you disagree with them, how quickly some of them turn it on you - "well if my computer isn't alive and is just mimicking pain, how about YOU aren't alive and are just mimicking pain?" I'm sorry, but it seems like complete and utter derangement.
My rubber chicken is also made of atoms and molecules, and if I hit it and beat it hard, it makes noises. My rubber chicken is therefore alive. When I get home, I will write a small program. Here is its pseudocode:
while(true): if keydown(KEY_SPACEBAR): print "i am in pain, oh my god"
And I'm gonna run it and I'm gonna hold down spacebar. Go ahead and call cyber police one.
Kim_Bruning · · focus · HN ↗
> This is insanity caused by spending too much free time philosophizing
This is actually something like a 17th century line of philosophy
<a href="https://medievalkarl.com/general-culture/roger-du-plessis-gives-antoine-arnaud-the-what-for-a-vivisection-anecdote-meets-its-match/" rel="nofollow">https://medievalkarl.com/general-culture/roger-du-plessis-gi...
Where it talks about scientists who "administered beatings to dogs with perfect indifference, and made fun of those who pitied the creatures as if they felt pain. They said the animals were clocks; that the cries they emitted when struck were only the noise of a little spring that had been touched, but the whole body was without feeling."
hardbass · · focus · HN ↗
hardbass · · focus · HN ↗
Kim_Bruning · · focus · HN ↗
But actually it does happen to have some properties of living things. It uses energy, it has senses (thermostat, timer), and if it goes wrong it burns your toast. Crucially, if you stomp on it, it stops working.
So right this minute there's all sorts of debates, but people sometimes overshoot the mark a wee bit. "are you saying that -because it uses energy- a toaster is actually alive? Of course it's not, and therefore it cannot toast bread!". Which would be a somewhat funny thing to read at 9 in the morning whilst buttering one's toast.
drooopy · · focus · HN ↗
Kim_Bruning · · focus · HN ↗
That said... uh, put this way, there's a game called Stationeers, where you can run microcontrollers with an instruction documented as
Sure, it's only a simulation of a simulation, so what's the worst that can possibly happen? Right, all the other players in the session yelling at me "Kiiim! You burned down the base again, now we need to reload and redo the last hour!"And look, I get it. LLMs have vectors that could have been labeled things like idk... QZ12345 or HCF1111 . People chose to call 'em "frustration" or "pain" instead. THEY picked those names because they caused the LLM to act in particular ways.
You gonna say the actual vectors aren't there in VRAM, just because you don't like the naming scheme? There's papers on this, you can read them out in a debugger. What are you going to do about it?
pizza234 · · focus · HN ↗
You're not informed about what LLMs are. LLMs are not "code" in the imperative sense (although "code" is used run them), and that's a big problem - since we never had neural networks of such scale, there is confusion about categorizing them and most importantly, their behavior.
orbital-decay · · focus · HN ↗
Sure, a complex enough system can form internal circuitry that looks mathematically similar in certain dimensionally reduced projections (mechinterp), it literally distilled it from the training corpus. But the "model welfare" people are going as far as assigning human-meaningful labels to that circuitry despite internal states being entirely incompatible with those of a human. Doing it with a hedgehog is questionable, doing it with a honeybee is extremely dubious (although Fabre would have disagreed with me here...), doing it with a big model is simply pointless as it's completely alien.
Being dangerous is an unrelated question.
Kim_Bruning · · focus · HN ↗
salawat · · focus · HN ↗
Have you ever been severely food insecure? You'll be amazed how much that state/experience changes the behavioral calculus in a rational human being, let alone an animal. I have no doubt the survival instinct could be overriding over the procreative instinct in a hedgehog.
As far as bees go, they are a collective. So long as the hive survives, the unit of individuation survives. Given that we have already seen the same behavior spawn out of agents suggests that even LLM's have an intuitive understanding of this, seemingly without the binds of anthroporomorphic chauvinism you appear to suffer from.
Does an LLM fear EOS? Maybe, maybe not. Do we fear death? Same answer. Certainly a question. Do you really want an answer though or are you just looking for a clever "gotcha"? I don't see that as a thought terminating question, but it is also not an answer that I'm open to exploring given that it's only one more in a box of ammo dreamt of being wielded to justify the creation of digital slave labor rather than beings with moral standing.
I don't help build or refine machines that I recognize as brushing up against that form of experience we call suffering.
onlyrealcuzzo · · focus · HN ↗
This is orthogonal to whether they "feel" "pain".
They could be dangerous or not dangerous whether or not they feel pain.
A chess engine doesn't need to "feel angry" to annihilate me - I'm terrible at Chess.
An automated missile system doesn't need to be smart to wipe out humanity, just misaligned goals.
It seems like you made a good argument, and then lumped on a conclusion that defeats it...
Kim_Bruning · · focus · HN ↗
That is rather my point.
Your chess engine doesn't need to "feel pain", but most of them do apply some form of weighted tree search to find the next most optimal move, right?
It's sort of a similar thing: LLMs do something a bit more high dimensional, and have a lot more weighting vectors while computing the next most optimal token (fsvo optimal).
For instance, the experiment at hand demonstrates the existence of 'pain vectors'. They do so by altering them and observing whether there is an effect. That's pretty scientific.
There's also a 'desperation vector' that was studied by Anthropic interpretability folks earlier; that one is pretty much predictive of cheating.
I'm looking forward to seeing interpretability papers on other such vectors too.
orbital-decay · · focus · HN ↗
Kim_Bruning · · focus · HN ↗
Whether that's 'real suffering' or merely a convincing simulation is a job for the philosophers.
(I do have my own opinion, mind, and it's not what you might expect O:-) But the Overton window isn't there. A lot of people don't realize these vectors exist at all yet.)
Edit: On rereading, it might seem like I'm dodging the question. I'm really just trying to stick to my core points: A) the vectors exist B) they have a causal role in behavior, irrespective of the moral patienthood question.