This is so transparent. Their next move: 'during our testing GLM-5.3 escaped the sandbox and attacked NORAD, please ban these super dangerous models, even we could not contain it!'.
I wonder who the real audience of these messages is.
I seriously thought that's what they were going to say. It seemed like it was setting up the reader to think the agent was about to use the exploit to escape the sandbox. Then, it was contained by Anthropic's security practices. But, if random people run things things, they couldn't be contained! They'll be a hoard of agents attacking the whole Internet!
Then, it was just that it made an exploit, and works really well, and people might use it instead of their products. Tragic for their investors I guess...
jacquesm · · focus · HN ↗
I wonder who the real audience of these messages is.
nickpsecurity · · focus · HN ↗
Then, it was just that it made an exploit, and works really well, and people might use it instead of their products. Tragic for their investors I guess...