Several vulnerabilities have been discovered in the Linux kernel
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Several vulnerabilities have been discovered in the Linux kernel
Unofficial Hacker News client; not affiliated with Y Combinator.
kalessin · · focus · HN ↗
sashank_1509 · · focus · HN ↗
1. LLMs have high false positive rate. From mythos 79 vulnerabilities found in the Linux kernel, only a single digit were actual bugs and they were all obscure so don’t panic.
2. What does obscure mean? I don’t really understand it, but many of the bugs have to do with custom network drivers or other custom drivers that are very specific to certain organizational setups, not a general Linux distro issue.
3. He’s very frustrated with the high false positive rate mythos generates. Even after multiple rounds of adversarial review and prompting strats, he mentions it is > 20% false positive rate, which wastes a lot of time. When some random user on the internet brings up a bug with an LLM it’s almost always fake, he even says just push back a few times claiming it’s not a bug to see if it’s a real bug (LLMs very quickly cave and “notice their mistake” etc)
4. General observation on the useful bugs mythos finds. Chain multiple smaller bugs to see if you can get a bigger breakage. Mythos is really good at constructing these long convoluted chains that fuzzers miss.
5. Go through recent bug fixes and check if similar bugs are hidden elsewhere in the codebase. Mythos is good at such pattern matching albeit with a high false positive rate.
Final conclusion: don’t panic, the bugs are getting fixed, this is not as bad as the first fuzzer bug mania and will be fixed quicker, he estimates a year and we won’t see huge bug reports anymore.
pixl97 · · focus · HN ↗
If it's just 20% that is really low. Especially for complicated long chain potential bugs.
Most other detection tools have much higher rates of FP, or much higher rates of false negative.
Then you have humans that miss bugs for 20+ years. Or, they don't tell you about the things they thought were bugs they wasted hours on themselves. Because of this it's really hard to measure how bad/good the AI really is.
It would be interesting to know why the more SOTA models are getting the FPs. Is it from a lack of understanding of C? Is it complex code with deep branches? Is it code smell and convoluted logic?
p-o · · focus · HN ↗
keeda · · focus · HN ↗
If there's something wrong or prone to misinterpretation in the TL;DR it would be better to call it out in response to that, rather than the users responding to the TL;DR, simply because it's likely that's what most people will respond to.
voakbasda · · focus · HN ↗
sashank_1509 · · focus · HN ↗
pixl97 · · focus · HN ↗