‹ BackHN Continuity

Thread

The Download: why AI's latest breakthroughs and fears may be more hype than rea

49 points · 128 comments · joozio

  1. bananaflag · · focus · HN ↗
    > Describing them as “superintelligence” or “rogue models” ascribes agency to products rather than to the companies building them. This framing markets these companies’ products as “superhuman” and, at the same time, helps the companies evade accountability for their actions.

    If someone discovered how to summon demons to aid them in robbing banks the main issue wouldn't be "but who has the responsibility for the crime, the human or the demon", it would be "OMG DEMONS".

    Seriously, I don't get what sort of world these people are living in.

    1. rubendev · · focus · HN ↗
      That's not what the LLM hacking accidents have been like at all though. To improve the analogy, it would be like summoning a totally passive demon, giving it weapons and placing it next to a bank, place a small fence around it, and then command it to perform a totally safe "exercise" that is exactly like robbing a real bank.

      LLMs do not have agency, they are just producing tokens based on a prompt that a person entered, and some of these tokens can trigger the tools that a person gave them access to.

      1. 0xDEAFBEAD · · focus · HN ↗
        Dwarkesh, for one, defended his use of "anthropomorphic" language.

        >Sacrificing now yields Oracle for team, but forfeits our chance, question mark. But other agents were pushing it, sending a message saying, go, sacrifice final now. And then EarlyBig eventually agreed, thinking to itself, our own utility may be already near zero. Sacrifice rational.

        <a href="https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=X50zezLFWWI#t=2m" rel="nofollow">https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=X50zezLFWWI#t=2m

        My suspicion is that many of the &quot;LLMs do not have agency&quot; folks just haven&#x27;t learned much about the details of the incident. It was specifically with LLM agents that were trained to be more persistent than usual.

        If you&#x27;re going to say that the incident details don&#x27;t matter, and LLMs lack agency because it&#x27;s all based on floating-point math--why can&#x27;t I say that humans lack agency, because it&#x27;s all based on neurons firing?

        1. embedding-shape · · focus · HN ↗
          &gt; If you&#x27;re going to say that the incident details don&#x27;t matter, and LLMs lack agency because it&#x27;s all based on floating-point math--why can&#x27;t I say that humans lack agency, because it&#x27;s all based on neurons firing?

          They&#x27;re not saying this, we&#x27;re saying LLMs lack agency because if you run a LLM and don&#x27;t send any prompts, literally nothing happens.

          Instruct it to &quot;Find the right answer regardless of where&quot;, it&#x27;ll do exactly this. They&#x27;re passive in that they don&#x27;t act by themselves, somewhere, at one point, someone &quot;told&quot; the LLM to &quot;do something&quot; and that&#x27;s the cause and the reason for saying &quot;LLMs do not have agency&quot;.

          1. 0xDEAFBEAD · · focus · HN ↗
            Imagine a super-competent Navy SEAL who just sleeps in the barracks unless his commander tells him to do something. Does the Navy SEAL lack agency? As a target of this Navy SEAL, should you be reassured by the fact that they&#x27;ll be sleeping in the barracks unless their commander tells them to do something?
            1. diidjdicksodksk · · focus · HN ↗
              No, the Navy SEAL does not lack agency in this scenario because they are still a person with individual agency, they can act on their own accord without anyone prompting them to, but in this case their superior instructed them to stay put. He could rebel and go rogue, but that would mean he used his own personal agency to defy orders given to him, which again places responsibility on the actor that actually possesses agency.
              1. 0xDEAFBEAD · · focus · HN ↗
                Suppose this SEAL is very obedient by disposition so the probability of him going rogue is akin to the probability of an LLM hallucinating or whatever.
                1. tavavex · · focus · HN ↗
                  No matter how you slice it, it&#x27;s not possible to equate humans to tools like this. If the SEAL receives an order to stand at attention until told otherwise, he will not stay in place for weeks until starving. An algorithm will work towards its self-destruction if ordered to - it doesn&#x27;t care, it doesn&#x27;t have the biological signals telling it otherwise. If the SEAL is lost somewhere with no one to give him orders, he will soon start acting on his own to ensure his survival and comfort. A tool will not do anything until directly activated and used by something that does have an active will, like a human. If the SEAL is ordered to massacre his hometown, no matter how obedient he is, he might have reservations. An algorithm with no instincts to get in the way will do anything if this behavior isn&#x27;t somehow inhibited by people during training. It wouldn&#x27;t even need the explicit order, if an irresponsible operator tells it to accomplish something by any means, that means that anything is on the table. Good thing the AI labs aren&#x27;t stuffed to the brim with irresponsible operators.
                  1. 0xDEAFBEAD · · focus · HN ↗
                    What if I prompt the AI to behave like the SEAL in the relevant ways, and fine-tune until it does so?
                    1. tavavex · · focus · HN ↗
                      Then these actions are downstream of your will, not the AIs. It doesn&#x27;t change anything, you&#x27;d just be playing with an imitation. Ordering it to pretend to be like a human does not a human equivalent make.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.