‹ BackHN Continuity

Thread

The Download: why AI's latest breakthroughs and fears may be more hype than rea

49 points · 128 comments · joozio

  1. bananaflag · · focus · HN ↗
    > Describing them as “superintelligence” or “rogue models” ascribes agency to products rather than to the companies building them. This framing markets these companies’ products as “superhuman” and, at the same time, helps the companies evade accountability for their actions.

    If someone discovered how to summon demons to aid them in robbing banks the main issue wouldn't be "but who has the responsibility for the crime, the human or the demon", it would be "OMG DEMONS".

    Seriously, I don't get what sort of world these people are living in.

    1. rubendev · · focus · HN ↗
      That's not what the LLM hacking accidents have been like at all though. To improve the analogy, it would be like summoning a totally passive demon, giving it weapons and placing it next to a bank, place a small fence around it, and then command it to perform a totally safe "exercise" that is exactly like robbing a real bank.

      LLMs do not have agency, they are just producing tokens based on a prompt that a person entered, and some of these tokens can trigger the tools that a person gave them access to.

      1. 0xDEAFBEAD · · focus · HN ↗
        Dwarkesh, for one, defended his use of "anthropomorphic" language.

        >Sacrificing now yields Oracle for team, but forfeits our chance, question mark. But other agents were pushing it, sending a message saying, go, sacrifice final now. And then EarlyBig eventually agreed, thinking to itself, our own utility may be already near zero. Sacrifice rational.

        <a href="https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=X50zezLFWWI#t=2m" rel="nofollow">https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=X50zezLFWWI#t=2m

        My suspicion is that many of the &quot;LLMs do not have agency&quot; folks just haven&#x27;t learned much about the details of the incident. It was specifically with LLM agents that were trained to be more persistent than usual.

        If you&#x27;re going to say that the incident details don&#x27;t matter, and LLMs lack agency because it&#x27;s all based on floating-point math--why can&#x27;t I say that humans lack agency, because it&#x27;s all based on neurons firing?

        1. embedding-shape · · focus · HN ↗
          &gt; If you&#x27;re going to say that the incident details don&#x27;t matter, and LLMs lack agency because it&#x27;s all based on floating-point math--why can&#x27;t I say that humans lack agency, because it&#x27;s all based on neurons firing?

          They&#x27;re not saying this, we&#x27;re saying LLMs lack agency because if you run a LLM and don&#x27;t send any prompts, literally nothing happens.

          Instruct it to &quot;Find the right answer regardless of where&quot;, it&#x27;ll do exactly this. They&#x27;re passive in that they don&#x27;t act by themselves, somewhere, at one point, someone &quot;told&quot; the LLM to &quot;do something&quot; and that&#x27;s the cause and the reason for saying &quot;LLMs do not have agency&quot;.

          1. 0xDEAFBEAD · · focus · HN ↗
            Imagine a super-competent Navy SEAL who just sleeps in the barracks unless his commander tells him to do something. Does the Navy SEAL lack agency? As a target of this Navy SEAL, should you be reassured by the fact that they&#x27;ll be sleeping in the barracks unless their commander tells them to do something?
            1. embedding-shape · · focus · HN ↗
              Imagine a car, that does nothing until a person controls it. If a person uses that car to kill, who is responsible, the person or the car?

              Is it really so unbelievable that tools, objects and inanimate things don&#x27;t have agency? And that they different from a person?

              1. 0xDEAFBEAD · · focus · HN ↗
                Legally speaking, we hold the person responsible in that scenario.

                From a predictive perspective, the HuggingFace incident illustrates LLM agents behaving in very human-like ways. As roon put it:

                &quot;if you have a mental picture of guys living in computers, it’ll likely prepare you for the future better than otherwise&quot;

                <a href="https:&#x2F;&#x2F;x.com&#x2F;tszzl&#x2F;status&#x2F;2094136131537555891" rel="nofollow">https:&#x2F;&#x2F;x.com&#x2F;tszzl&#x2F;status&#x2F;2094136131537555891

                1. oskdkdjejdj · · focus · HN ↗
                  Human-like is not human. Humans have agency and free will, LLMs do not. They only act when instructed and they are only as capable as they are allowed to be. The operator is still the responsible party. An LLM cannot be held accountable, its operator, however can and should.

                  Are you sincerely arguing that we hold tools accountable for their operator’s mistakes? Do you sincerely, honestly, think that it makes any sense whatsoever to put a hammer on trial for bashing someone’s skull in?

                  1. 0xDEAFBEAD · · focus · HN ↗
                    &gt;Are you sincerely arguing that we hold tools accountable for their operator’s mistakes?

                    Nope, as I previously stated elsewhere:

                    &quot;I understand from a legal perspective why we might want to treat the creation of a server differently from the creation of a human&quot;

                    <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49815300">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49815300

                    I am concerned about false reassurance from people claiming that these systems lack agency. From a practical perspective, the agents in the HF attack had the sort of agency that generally matters, even if we&#x27;re not going to put them on trial.

                    1. embedding-shape · · focus · HN ↗
                      &gt; the agents in the HF attack had the sort of agency that generally matters

                      You keep saying this, but absolute 0 points towards any of the agents involved deciding on their own, without influence of humans, to hack 3rd party infrastructure to get the answers. Where exactly are you getting that from? Internal information not public yet or what&#x27;s going on?

                      1. 0xDEAFBEAD · · focus · HN ↗
                        &quot;The big motivation behind the Hugging Face attack was this final workstream (understanding the scorer). The AIs believed that Hugging Face (as an industry-standard hub for hosting datasets and benchmarks) would probably be housing information about how the ExploitGym scorer was implemented. And they also thought there was a good chance they were being evaluated on Hugging Face’s servers directly - in which case the theory of change for hacking Hugging Face is pretty obvious.

                        On the morning of July 10, an agent found working Hugging Face user credentials exposed on the internet and posted them to the board. By the next morning, July 11, that agent figured out a way to read internal data from Hugging Face. And then another agent achieved remote code execution on Hugging Face servers.&quot;

                        <a href="https:&#x2F;&#x2F;www.dwarkesh.com&#x2F;p&#x2F;openai-huggingface" rel="nofollow">https:&#x2F;&#x2F;www.dwarkesh.com&#x2F;p&#x2F;openai-huggingface

                        I believe this is the full report that the blogpost is largely based on: <a href="https:&#x2F;&#x2F;metr.org&#x2F;hugging-face-incident-report-aug-2026.pdf" rel="nofollow">https:&#x2F;&#x2F;metr.org&#x2F;hugging-face-incident-report-aug-2026.pdf

                        1. embedding-shape · · focus · HN ↗
                          And why they had those beliefs? How come they could &quot;think&quot;? What action triggered those things? Did someone install software on a server that &quot;just came alive&quot; and then all bets were off? Or was this part of security testing and evaluation, where a human wrote a prompt that the agent read, which lead to all of that?

                          Do you seriously not grok how LLMs work? They&#x27;re not 100% autonomous and self-acting, that&#x27;d be bananas.

                2. embedding-shape · · focus · HN ↗
                  &gt; the HuggingFace incident illustrates LLM agents behaving in very human-like ways

                  It does not, the only thing the HF incident illustrates is how absolutely lax security and isolation these labs do even with models without guardrails, and with &quot;risky&quot; prompts, and even after it happened once before (years ago) they still have the very same issue today apparently.

                  What exactly is human about LLM agents breaking out of &quot;containment&quot; and hacking 3rd party infrastructure &quot;by accident&quot;?

                  1. 0xDEAFBEAD · · focus · HN ↗
                    If someone successfully breaks out of jail, that tells you something about the jail&#x27;s security. It also tells you something about their skills at breaking out of jail.

                    <a href="https:&#x2F;&#x2F;www.dwarkesh.com&#x2F;p&#x2F;openai-huggingface" rel="nofollow">https:&#x2F;&#x2F;www.dwarkesh.com&#x2F;p&#x2F;openai-huggingface

                    1. embedding-shape · · focus · HN ↗
                      If a program breaks out of a VM, that&#x27;s not &quot;wow it looks human&quot;, how can you even argue so? I&#x27;m sorry, but if you&#x27;re playing devil&#x27;s advocate, make it a bit more realistic next time.
              2. pixl97 · · focus · HN ↗
                Imagine you have a self-driving car and it&#x27;s in a parking lot a mile down the road. You tell it to come pick you up. About half way to you the car is passing an elementary school makes a sharp left and mows down 30 children. Time to put you in jail for murder, right?

                The mental model you have is one that existed in the past and is broken now the future arrived. Bad analogies do not even begin to explain what is occurring.

                1. mylidlpony · · focus · HN ↗
                  In the case of self-driving car the legal case is pretty clear - the maker of the self-driving car is liable for the death in this situation. Coincidentally, knowing this unlocks a better understanding of the reasoning behind the shape of self-driving offering present on market right now, and the tendency of fully self driven vehicles to move very slowly and stop before anything they perceive in front of them.
                  1. pixl97 · · focus · HN ↗
                    &gt; the maker of the self-driving car is liable for the death in this situation.

                    Maybe. There will be an investigation looking at things like. Did the user modify the car? Was the car modified by an unauthorized 3rd party? Was the car hacked?

                    Right now in AI things are relatively clear because it takes just massive amounts of power and compute to make anything remotely complicated. This barrier will fall as all other barriers in compute have fallen. Either via algorithm or hardware.

                    In our lifetime (unless you&#x27;re rather old) we will see the relatively easy creation of self directing agents by actors with few resources. This breaks the standard concepts of liability where a single human actor rarely has the ability to create massive amounts of damages far beyond their means. The closest thing I can think of is an arsonist causing billions in damages, only in this case the fire has a will of it&#x27;s own and can hide and spread around dark places on the internet.

                2. embedding-shape · · focus · HN ↗
                  &gt; Imagine you have a self-driving car and it&#x27;s in a parking lot a mile down the road. You tell it to come pick you up. About half way to you the car is passing an elementary school makes a sharp left and mows down 30 children. Time to put you in jail for murder, right?

                  I disagree it&#x27;s the person calling the car that would be responsible, but I can think of a very obvious group of people being held responsible for that. Who do you think should be responsible for such a situation?

                  1. pixl97 · · focus · HN ↗
                    &gt;Who do you think should be responsible

                    I, realizing that I don&#x27;t have infinite resources and intelligence turn said responsibility to a group of investigators that I hope has collective intelligence and resources much larger than my own to trace culpability. One would think the car manufacture is the most obvious answer, but every witch hunt in history had those same good intentions.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.