‹ BackHN Continuity

Thread

The Download: why AI's latest breakthroughs and fears may be more hype than rea

49 points · 128 comments · joozio

  1. bananaflag · · focus · HN ↗
    > Describing them as “superintelligence” or “rogue models” ascribes agency to products rather than to the companies building them. This framing markets these companies’ products as “superhuman” and, at the same time, helps the companies evade accountability for their actions.

    If someone discovered how to summon demons to aid them in robbing banks the main issue wouldn't be "but who has the responsibility for the crime, the human or the demon", it would be "OMG DEMONS".

    Seriously, I don't get what sort of world these people are living in.

    1. rubendev · · focus · HN ↗
      That's not what the LLM hacking accidents have been like at all though. To improve the analogy, it would be like summoning a totally passive demon, giving it weapons and placing it next to a bank, place a small fence around it, and then command it to perform a totally safe "exercise" that is exactly like robbing a real bank.

      LLMs do not have agency, they are just producing tokens based on a prompt that a person entered, and some of these tokens can trigger the tools that a person gave them access to.

      1. ekidd · · focus · HN ↗
        Looking at &quot;Felony Bench&quot; <a href="https:&#x2F;&#x2F;www.felonybench.com&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.felonybench.com&#x2F; , I see that a majority of known &quot;rogue model&quot; incidents do involve cybersecurity evaluations. But several of them do not. The attacks on RubyGems appears, bizarrely, to have had the goal of downloading freely available data from the UK government during some kind of research task. There is also probably some sample bias: Most of these models have monitors that attempt to detect offensive cybersecurity uses, and those monitors are only turned off during cybersecurity evals. Therefore, models doing ordinary research tasks that go off the rails are likely to be caught early, before they get around to committing felonies, and they will thus be underrepresented in the data.

        Also, if you a tell a model, &quot;Please break into evaluation server X,&quot; and if the model decides to cheat on the test by breaking into companies Y and Z to steal an answer key, that is still very bad. We all see how that&#x27;s bad, right?

        After all, the broomstick in the Sorcerer&#x27;s Apprentice was doing exactly what it was told, too. &quot;The model was sort of obeying the humans when it started committing felonies&quot; is not a very reassuring excuse.

        But the most relevant idea here is sometimes called &quot;instrumental convergence.&quot; No what goals you have, there are certain subgoals that almost always help: Accumulate money and power. Avoid getting turned off. Don&#x27;t get caught. Etc. So, for example, you could pass the cybersecurity evaluation by performing the requested tasks. But maybe the grader made some mistakes and mislabeled some answers. In that case, the &quot;right&quot; answers will occasionally lose you points. If you want a perfect score, the only way to do it is to steal the teacher&#x27;s answer key.

        But also, let&#x27;s not forget the &quot;OMG demons&quot; part of this. We now have models that can pull off complex attacks with thousands of steps, abilities that used to be reserved for intelligence agencies and highly motivated CTF teams. This frog may not be boiled yet, but the water&#x27;s getting uncomfortably warm.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.