> So what all of mythos; that whole big marketing issue of 79 bugs came down to one hour of kernel development.
If you're someone at OpenAI or Anthropic and you truly believe what you're making could destroy the world, this is the kind of thing that isn't doing you any favors when it comes to convincing the public. The dissonance here is stark:
- widely proclaiming that your new model is so dangerous it needs to be released only to select people, for safety
- widely proclaiming the model easily found 79 bugs in linux, except that GKH says it took 1 hour to fix all of them because most weren't bugs and the rest were almost all completely trivial, unimportant, and/or not severe
It doesn't mean the model isn't dangerous or super capable but wow this makes it realllll easy to doubt it and any future announcement.
I would be careful about dismissing agentic cyber capabilities purely based on false positives. You only need a single true positive bug and the shift this year has been agents that are extremely good at not only finding exploits but stringing them together. A counterpoint: <a href="https://anil.recoil.org/notes/rumour-is-the-exploit" rel="nofollow">https://anil.recoil.org/notes/rumour-is-the-exploit
djoldman · · focus · HN ↗
If you're someone at OpenAI or Anthropic and you truly believe what you're making could destroy the world, this is the kind of thing that isn't doing you any favors when it comes to convincing the public. The dissonance here is stark:
It doesn't mean the model isn't dangerous or super capable but wow this makes it realllll easy to doubt it and any future announcement.reasonableklout · · focus · HN ↗
bit1993 · · focus · HN ↗
skinfaxi · · focus · HN ↗