Revealing the details of how OpenAI agents hacked Hugging Face
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Revealing the details of how OpenAI agents hacked Hugging Face
Unofficial Hacker News client; not affiliated with Y Combinator.
grim_io · · focus · HN ↗
The frontier labs can monitor the behavior of agents for millions of customers (did you try hacking with frontier labs? Good luck), but they can't secure internal use?
Give me a break. What a bunch of amateurs.
stratos123 · · focus · HN ↗
They "monitor" this by having classifiers watching the model output that'd stop the session/punt you to a weaker model/raise an alarm if they see anything suspicious. They can't do that in a cybersec eval because the normal safeguards would just be going off at all times.
Why didn't they attach a special classifier, which'd allow hacking-within-the-task but not going off the rails? Good question; part of the answer is obviously "it's hard to have a classifier that smart" and "it'll have false positives" but even a very bad safeguard would have stopped this.
ses1984 · · focus · HN ↗