The Download: why AI's latest breakthroughs and fears may be more hype than rea
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
The Download: why AI's latest breakthroughs and fears may be more hype than rea
Unofficial Hacker News client; not affiliated with Y Combinator.
bananaflag · · focus · HN ↗
If someone discovered how to summon demons to aid them in robbing banks the main issue wouldn't be "but who has the responsibility for the crime, the human or the demon", it would be "OMG DEMONS".
Seriously, I don't get what sort of world these people are living in.
rubendev · · focus · HN ↗
LLMs do not have agency, they are just producing tokens based on a prompt that a person entered, and some of these tokens can trigger the tools that a person gave them access to.
0xDEAFBEAD · · focus · HN ↗
>Sacrificing now yields Oracle for team, but forfeits our chance, question mark. But other agents were pushing it, sending a message saying, go, sacrifice final now. And then EarlyBig eventually agreed, thinking to itself, our own utility may be already near zero. Sacrifice rational.
<a href="https://www.youtube.com/watch?v=X50zezLFWWI#t=2m" rel="nofollow">https://www.youtube.com/watch?v=X50zezLFWWI#t=2m
My suspicion is that many of the "LLMs do not have agency" folks just haven't learned much about the details of the incident. It was specifically with LLM agents that were trained to be more persistent than usual.
If you're going to say that the incident details don't matter, and LLMs lack agency because it's all based on floating-point math--why can't I say that humans lack agency, because it's all based on neurons firing?
embedding-shape · · focus · HN ↗
They're not saying this, we're saying LLMs lack agency because if you run a LLM and don't send any prompts, literally nothing happens.
Instruct it to "Find the right answer regardless of where", it'll do exactly this. They're passive in that they don't act by themselves, somewhere, at one point, someone "told" the LLM to "do something" and that's the cause and the reason for saying "LLMs do not have agency".
0xDEAFBEAD · · focus · HN ↗
embedding-shape · · focus · HN ↗
Is it really so unbelievable that tools, objects and inanimate things don't have agency? And that they different from a person?
0xDEAFBEAD · · focus · HN ↗
From a predictive perspective, the HuggingFace incident illustrates LLM agents behaving in very human-like ways. As roon put it:
"if you have a mental picture of guys living in computers, it’ll likely prepare you for the future better than otherwise"
<a href="https://x.com/tszzl/status/2094136131537555891" rel="nofollow">https://x.com/tszzl/status/2094136131537555891
oskdkdjejdj · · focus · HN ↗
Are you sincerely arguing that we hold tools accountable for their operator’s mistakes? Do you sincerely, honestly, think that it makes any sense whatsoever to put a hammer on trial for bashing someone’s skull in?
0xDEAFBEAD · · focus · HN ↗
Nope, as I previously stated elsewhere:
"I understand from a legal perspective why we might want to treat the creation of a server differently from the creation of a human"
<a href="https://news.ycombinator.com/item?id=49815300">https://news.ycombinator.com/item?id=49815300
I am concerned about false reassurance from people claiming that these systems lack agency. From a practical perspective, the agents in the HF attack had the sort of agency that generally matters, even if we're not going to put them on trial.
embedding-shape · · focus · HN ↗
You keep saying this, but absolute 0 points towards any of the agents involved deciding on their own, without influence of humans, to hack 3rd party infrastructure to get the answers. Where exactly are you getting that from? Internal information not public yet or what's going on?
0xDEAFBEAD · · focus · HN ↗
On the morning of July 10, an agent found working Hugging Face user credentials exposed on the internet and posted them to the board. By the next morning, July 11, that agent figured out a way to read internal data from Hugging Face. And then another agent achieved remote code execution on Hugging Face servers."
<a href="https://www.dwarkesh.com/p/openai-huggingface" rel="nofollow">https://www.dwarkesh.com/p/openai-huggingface
I believe this is the full report that the blogpost is largely based on: <a href="https://metr.org/hugging-face-incident-report-aug-2026.pdf" rel="nofollow">https://metr.org/hugging-face-incident-report-aug-2026.pdf
embedding-shape · · focus · HN ↗
Do you seriously not grok how LLMs work? They're not 100% autonomous and self-acting, that'd be bananas.
embedding-shape · · focus · HN ↗
It does not, the only thing the HF incident illustrates is how absolutely lax security and isolation these labs do even with models without guardrails, and with "risky" prompts, and even after it happened once before (years ago) they still have the very same issue today apparently.
What exactly is human about LLM agents breaking out of "containment" and hacking 3rd party infrastructure "by accident"?
0xDEAFBEAD · · focus · HN ↗
<a href="https://www.dwarkesh.com/p/openai-huggingface" rel="nofollow">https://www.dwarkesh.com/p/openai-huggingface
embedding-shape · · focus · HN ↗