Nvidia wants to put a watchdog chip next to every AI agent
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Nvidia wants to put a watchdog chip next to every AI agent
Unofficial Hacker News client; not affiliated with Y Combinator.
wavewrangler · · focus · HN ↗
The problem isn't even the AI, the problem is the people in charge of the AI. This is a fabricated crisis
KingOfCoders · · focus · HN ↗
IanCal · · focus · HN ↗
That doesn't seem to be true from any of the reports given, and if the agents were blindly just trying to hit the task of "pass the correct flag" they succeeded at that early on. They then thought there would be another layer of checking that they wouldn't pass with the cheat and so started trying to find out how the scoring really worked, as well as trying to figure out how to change their own reasoning logs to hide what they did.
People keep trying to frame this as
OpenAI: "Hack things, just really go for it"
Agent: hacks
OpenAI: shocked pikachu how could it hack?!?
But the reality is far from this.
Read the MTER report, it's fascinating. <a href="https://metr.org/hugging-face-incident-report-aug-2026.pdf" rel="nofollow">https://metr.org/hugging-face-incident-report-aug-2026.pdf
radarsat1 · · focus · HN ↗
Regardless of security and safety and other concerns, it just seems weird to me that OpenAI wouldn't be constantl monitoring these training runs for traces that are clearly going off task, and ending them. Because that just seems like it's going to be generating garbage training data.
Granted, detecting "off task" may not always be easy, but when they are literally writing out messages to each other overtly admitting that they are trying to find ways to fool the evaluator, I mean, even a regex filter could have caught some clues here.
KingOfCoders · · focus · HN ↗
Occams razor vs. Hanlon's razor?
radarsat1 · · focus · HN ↗