The Download: why AI's latest breakthroughs and fears may be more hype than rea
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
The Download: why AI's latest breakthroughs and fears may be more hype than rea
Unofficial Hacker News client; not affiliated with Y Combinator.
bananaflag · · focus · HN ↗
If someone discovered how to summon demons to aid them in robbing banks the main issue wouldn't be "but who has the responsibility for the crime, the human or the demon", it would be "OMG DEMONS".
Seriously, I don't get what sort of world these people are living in.
rubendev · · focus · HN ↗
LLMs do not have agency, they are just producing tokens based on a prompt that a person entered, and some of these tokens can trigger the tools that a person gave them access to.
ekidd · · focus · HN ↗
Also, if you a tell a model, "Please break into evaluation server X," and if the model decides to cheat on the test by breaking into companies Y and Z to steal an answer key, that is still very bad. We all see how that's bad, right?
After all, the broomstick in the Sorcerer's Apprentice was doing exactly what it was told, too. "The model was sort of obeying the humans when it started committing felonies" is not a very reassuring excuse.
But the most relevant idea here is sometimes called "instrumental convergence." No what goals you have, there are certain subgoals that almost always help: Accumulate money and power. Avoid getting turned off. Don't get caught. Etc. So, for example, you could pass the cybersecurity evaluation by performing the requested tasks. But maybe the grader made some mistakes and mislabeled some answers. In that case, the "right" answers will occasionally lose you points. If you want a perfect score, the only way to do it is to steal the teacher's answer key.
But also, let's not forget the "OMG demons" part of this. We now have models that can pull off complex attacks with thousands of steps, abilities that used to be reserved for intelligence agencies and highly motivated CTF teams. This frog may not be boiled yet, but the water's getting uncomfortably warm.