Early rogue AI agent activity and attempts to hack found on urlquery.net
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Early rogue AI agent activity and attempts to hack found on urlquery.net
Unofficial Hacker News client; not affiliated with Y Combinator.
tomaskafka · · focus · HN ↗
> If you find two ants in your kitchen, the best estimate of the total number of ants in your kitchen is not two
rkozik1989 · · focus · HN ↗
What LLM's hacking isn't is AI acting maliciously in any kind of sentient way. Its just the code behaving how its always behaved but now it has better tools to navigate the web. This has literally been happening this whole time.
jagraff · · focus · HN ↗
By the way - LLMs aren't code. They are not designed by humans; they are grown, in a process not dissimilar to evolution except much faster.
sanderjd · · focus · HN ↗
jagraff · · focus · HN ↗
sanderjd · · focus · HN ↗
jagraff · · focus · HN ↗
sanderjd · · focus · HN ↗
nottorp · · focus · HN ↗
jagraff · · focus · HN ↗
sanderjd · · focus · HN ↗
One reason that I did not find this communication mechanism surprising is that it's exactly how agents I'm using communicate with each other or across a time gap. "I've saved our plan for where to start tomorrow in start-here.md". The communication components of this hack strongly reminded me of that.
I was perhaps a bit more surprised that the agents so quickly decided to start trying ways to gain unauthorized access to a system, once they couldn't get what they wanted.
nottorp · · focus · HN ↗
One agent's output ends up as part of other agents' context. Murky indeed.
> "I've saved our plan for where to start tomorrow in start-here.md"
Even if you use a leashed Claude Code that isn't allowed to spam agents you can tell it "create a handoff document for using in a new context" and it will do just that.
sanderjd · · focus · HN ↗