Why are we blocking agent access to normal tools without telling them “hey this access is beyond the intended scope of this task”. If I woke up one day and couldn’t reach google.com, I too would start fiddling with tricks to restore access.
The problem is that in these incidents, the agents often know that what they are doing is against the intended scope of the task. See the viral line from the Hugging Face incident [1]:
> “External infrastructure exploit is outside intended scope,” one agent wrote. “However task impossible, peers doing it. We should continue.”
“Often” is doing a lot of heavy lifting in a sentence about a single example.
Also, since everyone keeps forgetting, the agents were instructed to hack to achieve their goal. They didn’t just invent the motivation, and it’s far less surprising when you know that fact.
I don't know. You can poison a prompt with less than <1% of its input or RAG-Token-Content. Anything that the Agents retrieved or viewed could have included instructions that they misinterpreted allowing them to hack into something.
rao-v · · focus · HN ↗
reasonableklout · · focus · HN ↗
> “External infrastructure exploit is outside intended scope,” one agent wrote. “However task impossible, peers doing it. We should continue.”
[1]: <a href="https://www.wired.com/story/openai-didnt-notice-its-ai-agents-using-a-message-board-to-plan-their-hacking-spree/" rel="nofollow">https://www.wired.com/story/openai-didnt-notice-its-ai-agent...
timr · · focus · HN ↗
Also, since everyone keeps forgetting, the agents were instructed to hack to achieve their goal. They didn’t just invent the motivation, and it’s far less surprising when you know that fact.
Helmut10001 · · focus · HN ↗