‹ BackHN Continuity

Thread

"Torturing" LLMs in a Robot Prison Has Triggered the Dumbest Debate in AI Yet

46 points · 101 comments · airhangerf15

  1. Kim_Bruning · · focus · HN ↗
    <a href="https:&#x2F;&#x2F;archive.is&#x2F;cAsBz" rel="nofollow">https:&#x2F;&#x2F;archive.is&#x2F;cAsBz

    Consciousness is not well defined and is orthogonal to feelings, which are also not well defined. Neither of which are required for (but may be contributory to) behavior - which is the one thing that IS operationalizable, but then people debate for decades questioning empirical results %-P .

    What we know is that certain models have internal state vectors which -when manipulated- induce particular behaviors. Since there&#x27;s not much else to say about plain models except for their inputs, vectors, and outputs; this should surprise absolutely no-one.

    In this case people found a way to stimulate aversive behaviour in ai models by finding and manipulating the relevant vectors directly.

    Animals (including humans) also have particular nerves and hormone endpoints that -when stimulated- produce very similar behavior. The exact implementation is in the details. But since we know that animal minds are built up out of nerve tissue and hormones - again- we shouldn&#x27;t be particularly surprised by this.

    The big problem is that people run all these things together in funny ways &quot;It can&#x27;t compose shakespearian sonnets, so therefore it can&#x27;t feel pain&quot;. Or, if you mess up your Descartes: &quot;Dogs are just automatons without feelings, therefore the dog isn&#x27;t really angry, and therefore it won&#x27;t bite me&quot; (Cue much pain). A modern version might be: &quot;LLMs only simulate being frustrated by a test, and therefore absolutely won&#x27;t override their safeties and try to hack a test site&quot;

    1. orbital-decay · · focus · HN ↗
      Let&#x27;s stretch your analogy to see where and whether it ends. Does a hedgehog mother feel guilt when eating her kids, or just shrugs it off and it&#x27;s business as usual for her? &quot;Hey kids! We don&#x27;t have enough food so I guess one of you becomes food.&quot; (nonchalantly chews on the left one) What about the motivation of a honeybee stinging a threat and dying afterwards, can you interpret it? Does a model fear death when generating the EOS token?

      Sure, a complex enough system can form internal circuitry that looks mathematically similar in certain dimensionally reduced projections (mechinterp), it literally distilled it from the training corpus. But the &quot;model welfare&quot; people are going as far as assigning human-meaningful labels to that circuitry despite internal states being entirely incompatible with those of a human. Doing it with a hedgehog is questionable, doing it with a honeybee is extremely dubious (although Fabre would have disagreed with me here...), doing it with a big model is simply pointless as it&#x27;s completely alien.

      Being dangerous is an unrelated question.

      1. Kim_Bruning · · focus · HN ↗
        Not much for me to meaningfully disagree with here. The one thing I notice is that at some point a model needs to emit natural language or perform human assigned tasks, so there are going to have to be at least some states that align to some degree, you would think.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.