‹ BackHN Continuity

Thread

"Torturing" LLMs in a Robot Prison Has Triggered the Dumbest Debate in AI Yet

46 points · 101 comments · airhangerf15

  1. Svip · · focus · HN ↗
    If we talk about pain specifically, it is will established that physical and phycological pain can be observed on instruments; stress produces chemical reactions, that can make needles flicker on meters. Even plants can feel pain in an observable manner.

    As far as I am aware, the GPUs don't work harder, don't heat up more, and nothing in their mathematical equations are different, when the result of these vector calculations produce text, that to humans look like someone in pain. Whereas real pain is objectively observable, LLMs' pain - so far at least - is entirely subjectively observable.

    That being said, I'm a fan of Star Trek's depiction of Data in TNG and the Doctor in VOY fighting for their rights to be treated as equals to the biological crew; and while I support their plight, as far as I can remember, the shows never made a claim that their pain was observable by instruments, thus making the judgement an entirely subjective decision (particularly in the case of the Doctor). Though I'd be glad to be corrected in this matter.

    1. bondarchuk · · focus · HN ↗
      >nothing in their mathematical equations are different, when the result of these vector calculations produce text, that to humans look like someone in pain.

      So at least something is different, namely the text output changes (and therefore some internal state too, of course). I think your analogy is too simple, it does not have to be the case that the GPU has to act like the equivalent of the body for an LLM in this respect.

      1. danaris · · focus · HN ↗
        But this is basically just

        Human: "Act like you're in pain."

        Computer: "Aaugh! It hurts! Why?! No more!"

        Human: "OMG! The computer feels pain!!"

        1. Kim_Bruning · · focus · HN ↗
          Almost.

          Human: "Let me just poke this vector and see what happens"

          LLM: "OUW!"

          The underlying experiment didn't tell the LLM what to do. Instead, the experimenters modified a vector and observed the outcome; thus showing that there is a vector that makes an LLM go "Ouw" . People called it a "pain vector", because that's easier to remember than , idk, LVF12345.

          1. techjamie · · focus · HN ↗
            Somewhere in an LLM are a bunch of vectors you can tweak to do anything. If you knew the right numbers to tweak, you could make the LLM always talk like Shakespeare.

            People and animals have emotions because they were developed under evolutionary pressure that made emotional animals better fit to survive. An animal that can feel anger or fear is more fit to survive than one that doesn't. But LLMs aren't put under those same pressures. Their evolutionary pressure is to be a good text predictor.

            But we can't subject them to the type of pain signal they experience during inference, because to the LLM, whether it's acting happy or pained, it's merely outputting what it's trained to be the most likely text to follow what came before.

            Edit: s/are/aren't/

            1. Kim_Bruning · · focus · HN ↗
              Right, normally "it's merely outputting what it's trained to be the most likely text to follow what came before."

              Now -while it's doing that- if you poke at certain vectors, its predictions will veer off course in interesting ways.

              This shows that the activation vectors exist, and that their modification provides a causal contribution to the output.

              Sure there's interesting consequences of that. But if we say that's the take-home message, that's good enough for me for one day!

              1. abandonliberty · · focus · HN ↗
                I asked AI to explain activation vectors. How much did it get wrong?
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.