"Sonnet 5.5’s cyber capabilities are a large improvement over Sonnet 5’s, so we’re deploying it with safeguards similar to those on Opus 5.5. Users can still find and fix bugs in their code as part of routine software development, but higher-risk cybersecurity tasks will visibly fall back to Sonnet 5
Sounds like at least for Anthropic models we reached peak cyber capabilities with Opus 4.8. Everything after that falls back to worse models
Hold on, is there any bar to begin with? For OpenAI's Daybreak Blue, I only had to go through the Persona KYC to gain access. With Anthropic's I had to submit links to my profile and briefly describe my use cases, which I doubt were read by any human being but at least there's some semblance of barrier.
How well it is documented or known that how they use the passport information and so on. Current blocker for EU citizen is to share that data for AI company…
Project Glasswing is by invitation only. <a href="https://www.anthropic.com/glasswing" rel="nofollow">https://www.anthropic.com/glasswing
Between OpenCode and Openrouter which one would you suggest? Sometimes I have pure grunt work to be done on non-sensitive data for which I want to use Chinese models. For example, tasks like extracting something from publicly available large pdf files.
opencode with model inference on cheaperinference.com has been working well for me - glm-5.3-flash is shockingly cheap (i've spent a total of a few dollars over several weeks of heavy usage), fast and capable for cyber tasks
The fun part is that the cash grab the frontier labs are running on cyber tasks might motivate enough people to pay for third parties; i.e it might bring enough cash to sustain Chinese competitors (and their open weights marketing strategy, which we all benefit from).
Where this gets painful, is that I was working on my own OSS project today, and this is what happened;
1. Opus 5.5 noted some concerns.
2. I asked it to write tests to safely test the concerns and write up a remediation plan
3. Opus 5.5 flagged and reset the conversation to Opus 4.8, murdering my usage quota and (possibly?) doing a sub-optimal set of tests.
NGL, it would have been at least polite for it to either:
1. Sanity check if I was the only/main committer on the repo (I'm the only committer, it's my side project) and then decide whether it was 'responsible' to help me fix it. (after all, I'm wanting to correct the problem, not exploit it!)
2. Warn me before just YOLOing the context to another model leading me to have to do a grace reset.
wongarsu · · focus · HN ↗
Sounds like at least for Anthropic models we reached peak cyber capabilities with Opus 4.8. Everything after that falls back to worse models
ttul · · focus · HN ↗
watusername · · focus · HN ↗
ttul · · focus · HN ↗
Daybreak Blue is the not the same thing as Daybreak Red, which has a more significant hurdle. I don't know anyone who has gotten access to Red.
nicce · · focus · HN ↗
stavros · · focus · HN ↗
ttul · · focus · HN ↗
SirSavary · · focus · HN ↗
stavros · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
gozzoo · · focus · HN ↗
Iolaum · · focus · HN ↗
malshe · · focus · HN ↗
Flere-Imsaho · · focus · HN ↗
malshe · · focus · HN ↗
beveradb · · focus · HN ↗
asp_hornet · · focus · HN ↗
I use GLM directly from z.ai, they do not retain or train on your data accordingly to their TOS.
Aissen · · focus · HN ↗
to11mtm · · focus · HN ↗
1. Opus 5.5 noted some concerns.
2. I asked it to write tests to safely test the concerns and write up a remediation plan
3. Opus 5.5 flagged and reset the conversation to Opus 4.8, murdering my usage quota and (possibly?) doing a sub-optimal set of tests.
NGL, it would have been at least polite for it to either:
1. Sanity check if I was the only/main committer on the repo (I'm the only committer, it's my side project) and then decide whether it was 'responsible' to help me fix it. (after all, I'm wanting to correct the problem, not exploit it!)
2. Warn me before just YOLOing the context to another model leading me to have to do a grace reset.