> So what all of mythos; that whole big marketing issue of 79 bugs came down to one hour of kernel development.
If you're someone at OpenAI or Anthropic and you truly believe what you're making could destroy the world, this is the kind of thing that isn't doing you any favors when it comes to convincing the public. The dissonance here is stark:
- widely proclaiming that your new model is so dangerous it needs to be released only to select people, for safety
- widely proclaiming the model easily found 79 bugs in linux, except that GKH says it took 1 hour to fix all of them because most weren't bugs and the rest were almost all completely trivial, unimportant, and/or not severe
It doesn't mean the model isn't dangerous or super capable but wow this makes it realllll easy to doubt it and any future announcement.
yes and it makes me wonder about the claims others make about their own experiences with the models too. is this the case of OAI lying, or is it part of ai psychosis where you literally lose touch with reality as you uncritically accept whatever claims the LLMs make?
> I don't understand the relevance of time to fix. It has no correlation with severity.
That they're small bugs that don't involve any serious architecture changes (or architecture sleuthing), just small pattern recognition. LLMs are pretty good (maybe great) at that. Finding 10 bugs is nice. Now find 10 more.
In the context of the talk, its about the message of "Don't Panic" because in reality none of this is nearly as bad as some people are making it out to be
I think the top-level comment is slightly missing the point of what GKH is saying: it's not 'it only took an hour to fix', it's "this amount of bugs is about what the kernel community finds and fixes each hour". i.e. this splashy announcement is really just a drop in the ocean of the volume that the kernel is handling.
It doesn't have to be world-ending. Autonomous, malicious agent swarms are something we haven't had to deal with. What does mitigation of a malicious, self-replicating swarm worm look like? We will find out soon enough.
Replication seems like it would be unlikely to matter, at least with the current trajectory. At the moment any models capable of this are really heavy, there's a limited number of places that they could replicate to and they will not at all be stealthy about it.
my idea of a self replicating worm would not just involve heavy models. It would be a mainly small model with enough instructions to spread and use/jailbreak available models to create a reasonably heavy one that can function without restrictions. So it would essentially prompted itself into existence.
I don't how feasible that is but with recent news of models communicating with each other/writing notes for itself. I think the idea is grounded enough for bigger models to write instructions or hid tools for smaller ones to use. Even updating the smaller models to behave differently.
> What does mitigation of a malicious, self-replicating swarm worm look like?
Much like the mitigation of Morris or Slammer. Self-replication is -in fact- an essential part of what makes a program a worm, the first of which was built and released in the early 1970s.
I would be careful about dismissing agentic cyber capabilities purely based on false positives. You only need a single true positive bug and the shift this year has been agents that are extremely good at not only finding exploits but stringing them together. A counterpoint: <a href="https://anil.recoil.org/notes/rumour-is-the-exploit" rel="nofollow">https://anil.recoil.org/notes/rumour-is-the-exploit
That's a different kind of capability though. None of the bugs discovered seem to be particularly serious and this talk pretty much shreds the idea that they're any good for that, but the thing called out here is that AI is very good at shortening the time to exploit security vulnerabilities. That's the thing that's really changed with threat management
djoldman · · focus · HN ↗
If you're someone at OpenAI or Anthropic and you truly believe what you're making could destroy the world, this is the kind of thing that isn't doing you any favors when it comes to convincing the public. The dissonance here is stark:
It doesn't mean the model isn't dangerous or super capable but wow this makes it realllll easy to doubt it and any future announcement.slopinthebag · · focus · HN ↗
jonahx · · focus · HN ↗
The headline here is that none of the bugs were serious.
pessimizer · · focus · HN ↗
That they're small bugs that don't involve any serious architecture changes (or architecture sleuthing), just small pattern recognition. LLMs are pretty good (maybe great) at that. Finding 10 bugs is nice. Now find 10 more.
20k · · focus · HN ↗
rcxdude · · focus · HN ↗
goolz · · focus · HN ↗
bauerd · · focus · HN ↗
rcxdude · · focus · HN ↗
riffruff24 · · focus · HN ↗
I don't how feasible that is but with recent news of models communicating with each other/writing notes for itself. I think the idea is grounded enough for bigger models to write instructions or hid tools for smaller ones to use. Even updating the smaller models to behave differently.
Gigachad · · focus · HN ↗
There isn’t too much of that sitting around unused right now.
bauerd · · focus · HN ↗
skinfaxi · · focus · HN ↗
simoncion · · focus · HN ↗
Much like the mitigation of Morris or Slammer. Self-replication is -in fact- an essential part of what makes a program a worm, the first of which was built and released in the early 1970s.
reasonableklout · · focus · HN ↗
bit1993 · · focus · HN ↗
skinfaxi · · focus · HN ↗
20k · · focus · HN ↗