"Sonnet 5.5’s cyber capabilities are a large improvement over Sonnet 5’s, so we’re deploying it with safeguards similar to those on Opus 5.5. Users can still find and fix bugs in their code as part of routine software development, but higher-risk cybersecurity tasks will visibly fall back to Sonnet 5
Sounds like at least for Anthropic models we reached peak cyber capabilities with Opus 4.8. Everything after that falls back to worse models
Hold on, is there any bar to begin with? For OpenAI's Daybreak Blue, I only had to go through the Persona KYC to gain access. With Anthropic's I had to submit links to my profile and briefly describe my use cases, which I doubt were read by any human being but at least there's some semblance of barrier.
Project Glasswing is by invitation only. <a href="https://www.anthropic.com/glasswing" rel="nofollow">https://www.anthropic.com/glasswing
wongarsu · · focus · HN ↗
Sounds like at least for Anthropic models we reached peak cyber capabilities with Opus 4.8. Everything after that falls back to worse models
ttul · · focus · HN ↗
watusername · · focus · HN ↗
stavros · · focus · HN ↗
ttul · · focus · HN ↗