‹ BackHN Continuity

Thread

"Torturing" LLMs in a Robot Prison Has Triggered the Dumbest Debate in AI Yet

46 points · 101 comments · airhangerf15

  1. Kim_Bruning · · focus · HN ↗
    <a href="https:&#x2F;&#x2F;archive.is&#x2F;cAsBz" rel="nofollow">https:&#x2F;&#x2F;archive.is&#x2F;cAsBz

    Consciousness is not well defined and is orthogonal to feelings, which are also not well defined. Neither of which are required for (but may be contributory to) behavior - which is the one thing that IS operationalizable, but then people debate for decades questioning empirical results %-P .

    What we know is that certain models have internal state vectors which -when manipulated- induce particular behaviors. Since there&#x27;s not much else to say about plain models except for their inputs, vectors, and outputs; this should surprise absolutely no-one.

    In this case people found a way to stimulate aversive behaviour in ai models by finding and manipulating the relevant vectors directly.

    Animals (including humans) also have particular nerves and hormone endpoints that -when stimulated- produce very similar behavior. The exact implementation is in the details. But since we know that animal minds are built up out of nerve tissue and hormones - again- we shouldn&#x27;t be particularly surprised by this.

    The big problem is that people run all these things together in funny ways &quot;It can&#x27;t compose shakespearian sonnets, so therefore it can&#x27;t feel pain&quot;. Or, if you mess up your Descartes: &quot;Dogs are just automatons without feelings, therefore the dog isn&#x27;t really angry, and therefore it won&#x27;t bite me&quot; (Cue much pain). A modern version might be: &quot;LLMs only simulate being frustrated by a test, and therefore absolutely won&#x27;t override their safeties and try to hack a test site&quot;

    1. onlyrealcuzzo · · focus · HN ↗
      &gt; LLMs only simulate being frustrated by a test, and therefore absolutely won&#x27;t override their safeties and try to hack a test site

      This is orthogonal to whether they &quot;feel&quot; &quot;pain&quot;.

      They could be dangerous or not dangerous whether or not they feel pain.

      A chess engine doesn&#x27;t need to &quot;feel angry&quot; to annihilate me - I&#x27;m terrible at Chess.

      An automated missile system doesn&#x27;t need to be smart to wipe out humanity, just misaligned goals.

      It seems like you made a good argument, and then lumped on a conclusion that defeats it...

      1. Kim_Bruning · · focus · HN ↗
        &gt; This is orthogonal to whether they &quot;feel&quot; &quot;pain&quot;. They could be dangerous or not dangerous whether or not they feel pain.

        That is rather my point.

        Your chess engine doesn&#x27;t need to &quot;feel pain&quot;, but most of them do apply some form of weighted tree search to find the next most optimal move, right?

        It&#x27;s sort of a similar thing: LLMs do something a bit more high dimensional, and have a lot more weighting vectors while computing the next most optimal token (fsvo optimal).

        For instance, the experiment at hand demonstrates the existence of &#x27;pain vectors&#x27;. They do so by altering them and observing whether there is an effect. That&#x27;s pretty scientific.

        There&#x27;s also a &#x27;desperation vector&#x27; that was studied by Anthropic interpretability folks earlier; that one is pretty much predictive of cheating.

        I&#x27;m looking forward to seeing interpretability papers on other such vectors too.

        1. orbital-decay · · focus · HN ↗
          Yeah but that&#x27;s not what model welfare (and the article) is about. It&#x27;s about people objecting to making the model suffer, quite literally.
          1. Kim_Bruning · · focus · HN ↗
            It&#x27;s about both, actually. The experimenters added or modified a &#x27;pain vector&#x27; in the model, and that&#x27;s how they&#x27;re &#x27;making it suffer&#x27; .

            Whether that&#x27;s &#x27;real suffering&#x27; or merely a convincing simulation is a job for the philosophers.

            (I do have my own opinion, mind, and it&#x27;s not what you might expect O:-) But the Overton window isn&#x27;t there. A lot of people don&#x27;t realize these vectors exist at all yet.)

            Edit: On rereading, it might seem like I&#x27;m dodging the question. I&#x27;m really just trying to stick to my core points: A) the vectors exist B) they have a causal role in behavior, irrespective of the moral patienthood question.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.