Greatly appreciated the candor. I've included a few slides into text that i thought were eye-opening to me:
From his Kernel Recipes 2026 slide on Mythos
```
Mythos's 79 vulnerabilities:
24 - no detail at all "something crashed"
14 - not a bug at all
3 - totally made up data
15 - already fixed in latest release
- 11 by others
- 4 by anthropic
20 - fixes were needed
- 7 "assume a malicious filesystem image"
- 2 "assume you can inject a malicious network packet into the middle of the stack"
- 2 "NOMMU"
- 6 sctp networking issues for untrusted devices
- 2 ipv6 minor network issues
- 1 gpu driver for local malicious user
```
GHK called this "10 'real' bugfixes", which to me sounds like there's a wild hype machine around these companies and uncritical parroting of every press release they make that falls apart when you engage the affected real experts.
We’ve seen this in a few open source repos we voluntarily manage security on. They’re not massive repos, but big enough they get attention from security researchers.
Most of the real low hanging fruit was picked up by humans years ago. When doing automated scanning, the majority of stuff is overly-verbose nonsense which takes hours of expert human labour to understand, test, and discard.
Reading through a Claude generated false positive is absolutely excruciating, because it is absolutely determined that what it’s found is justified. Often you’ll receive very long accompanying “proof of concept” code which demonstrates absolutely wild scenarios. It’s especially frustrating when you’re volunteering your time for a project, and a well-meaning contributor submits the report without the technical nous to understand why you’re rejecting it.
Okay so now they're like a top percentile fresh grad on meth. Still a lack of real world experience plus some bizarre failures that illustrate gaping holes in the mental model. Does that description work for you?
We've been seeing "But you're not using the latest model!" over and over again, with every new model supposedly "groundbreaking" and "a game-changer" - just for the general population to conclude a few months later that it once again doesn't live up to the crazy marketing hype.
Anthropic claimed that Mythos was so good at finding vulnerabilities that it was too dangerous to release to the public. As this post clearly shows: that (only again) simply isn't true. If you believe your favorite flavor of frontier model is the exception, it is up to you to provide proof to back up that claim.
Well let's see. You sell to management. Does management buy:
1) nuanced tools that talk back, question assumptions, take over decisions, ... oh and expose just how much management knows about the business. Or how little)
2) a tool that can provide the excuse "we've had our source checked and dealt with the remarks"
I wish people stopped equating LLM to interns or junior engineers. They are tools, and as good they may be at some specific things people do, they also really suck at many more which we wouldn’t find acceptable in humans.
It's merely a way to frame things, in terms of experience and trust. And it highlights how an LLM can code very well, but not truly understand the ramifications of that code.
An "eager, bright 20ish year old interns" will grow as you guide them, while LLM will not, they'll only grow when their owner (definitely not us) update them. So, trying to humanizing some tool with human emotion is wrong way to framing it. Treat tool as tool.
Indeed they are tools. But if people/companies treat them as tools that can take a problem a human used to solve, and make them solve it from start to end, then the comparison begins to be necessary.
> Right now, all top tier LLMs are as eager, bright 20ish year old interns.
> Yet directly guide them, monitor everything they do, some value emerges.
That's simply not true. I've had some very talented interns, and they are leagues ahead of what the LLMs can do. Not that it matters though, because the point of having interns wasn't to have them produce value. What made the investment worth it was that 12 months down the line I would have a competent colleague that I could have an interesting conversation with. A human person that could challenge some of my blind spots. A person that could take responsibility of something. Maybe not my most important work, but some of it. You don't get ANY of that from the LLM.
In some ways the reality is worse: the same version of an LLM won't get better at its job, even though you might get better at prompting it. Newer versions are trained on more code, which has obvious benefits, and their harnesses are better at taking advantage of tools that were created to keep human coders out of trouble.
The other side of the coin is that coding agents are not maximally productive unless you give them enough rope to potentially hang themselves. Over roughly the past year, the coding agents I use have gone from hot garbage to pretty consistently useful, especially if I find tasks where I can give them a lot of running room. On the other hand, last week I found a case where the coding agent was looping and flailing like it was doing every third try a year ago.
They fail less often, but they fail in the same way.
There is a vulnerability in library X when you call Y with these specially crafted parameters; as seen in the attached logs it can overflow buffer Z and clobber memory potentially leading to an RCE.
After:
Honest take: the load-bearing constraint violation is real. The log documenting the exploit gate is the ledger which weaves the story.
I think the way to handle this is to just feed it into another AI agent (a better one) and ask it how serious the issue actually is. Fight fire with fire!
What a great excerpt; thank you! It reminds me of what I find when I look at CVEs handed out by scanners at places I've worked for actual impact to systems I've owned... there are a lot of slop/false positives. (And that's even without "AI".)
That said, I remember trying to weigh the hype at the time of the announcement reading/skimming the papers Anthropic published, recognizing that bugcount alone wasn't super-relevant but also remember being impressed by an NFS bug and a kernel bug that struck me as relevant at the time. So where did that NFS issue show up in GKH's list you showed so nicely above?
It turns out, AFAICT, it's not on his list, but the reasons are perhaps interesting to others so I will post here. It turns out there were two NFS issues this past year conflated a bit in my memory:
* The Linux CVE-2026-31402 NFS heap overflow that could allow unauthenticated memory reads over the network isn't in that list of 79, presumably because it was found by Claude Code, not Mythos months earlier. (I am guessing it's not his "malicious network packet into the middle of the stack" and is a stronger attack being a remote attack.)
* And the CVE-2026-4747 NFS stack buffer overflow that allowed gaining full unauthenticated remote root access didn't show up in GKH's list of 79 because despite being Mythos-caught, it wasn't Linux, it was FreeBSD.
I guess this does match my memory now that I think about it, that there weren't any smoking Linux guns caught by Mythos.
* (I guess there was also a longstanding 27-year old OpenBSD TCP SACK-handling stack integer overflow than enabled remote crashes / Denial of Service found by Mythos.)
There is definitely Mythos hype, but just because it hit the BSD code base more than the GKH-managed Linux code base doesn't mean it was inappropriate to raise eyebrows from Mythos, in particular since "attacks only get better".
AI and Police have essentially the same journalists who, in lieu of any research or fact checking, just report verbatim their press releases and interviews.
(I think that's the Sherry Turkle episode I'm looking for)
the other I typically reference: <a href="https://www.npr.org/2025/07/18/g-s1177-78041/what-to-do-when-ai-says-i-love-you-we-talked-to-an-artificial-intimacy-expert" rel="nofollow">https://www.npr.org/2025/07/18/g-s1177-78041/what-to-do-when...
Mythos turned out to be exactly the marketing stunt it smelled like.
There are others like AISLE who seem to be a bit more successful in finding actual issues using LLMs in some shape or form though, whatever they do differently. Chances are high the secret sauce is not so much about the model being exceptionally powerful which would be bad news for the frontier labs.
here <a href="https://stanislavfort.substack.com/p/mythos-at-home-and-its-called-aisle" rel="nofollow">https://stanislavfort.substack.com/p/mythos-at-home-and-its-... they say "we match and beat Mythos, in some cases even using models you can run on your own hardware", so they seem to be using small models at least in some cases
I think Daniel is giving a more balanced view with:
> Any project that has not scanned their source code with AI powered tooling will likely find huge number of flaws, bugs and possible vulnerabilities with this new generation of tools. Mythos will, and so will many of the others.
Greg's video is a good reality check on the hype. But I'd be careful about generalizing from Linux, libcurl, etc which get far more scrutiny than software projects in general. LLM-assisted bug finding still matter a lot for everyday custom and less popular software.
people have different opinions of what Risk is. hence greg doesnt recognise certain things as risks that others do. That bein said most of those others wouldnt run Linux, and for sure 100% a shoe salesmen will try to sell you their shoes whatever the quality. They will also defend their price and quality whatever the quality though so that knife cuts both ways and on either end its consumer that gets cut
>7 "assume a malicious filesystem image"
If you ever used a USB storage device you're vulnerable to this one. Not even a strict chain of custody guarantees safety, because USB devices are often powered by exploitable programmable microcontrollers. If a known good USB device can be converted to a malicious USB device by unprivileged software, the malicious filesystem exploit becomes a local privilege escalation. It works better than tampering with the files on the filesystem because it escapes signature checks and gets you directly into kernel mode.
> ... or does the USB handshake happen with or involve the host OS?
It depends on how your host is configured: just as you can do GPU-passthrough, you can passthrough a single USB device or passthrough an entire USB controller to a VM.
Maybe I’m just naive but this breakdown signals a shocking lack of due diligence from anthropic. Did they put any effort into verifying these alleged bugs before sending them to maintainers?
The Linux Kernel is not necessarily the most interesting target for LLMs, since it gets a huge amount of attention. I would be interested to hear what people find in less prominent projects. For completeness that includes proprietary code stored in GitHub.
In our company we found real security issues in proprietary code using Opus 4.6/4.7. Obviously typical attackers might have difficulty finding these without code access but Claude was finding real CVEs, which we fixed.
p.s., This is an argument for not trying to deal with security problems by neutering the LLMs. To the extent LLMs are effective it weakens security.
And I suppose large companies like Microsoft, Google, Adobe, Apple (very famously sitting out the AI bubble) and Mozilla have been shipping record number of vulnerability fixes in their patches just because of the hype machine? ;-)
All of them are heavily invested in AI, and just because a change was merged doesn't mean it closed an actual vulnerability.
Let's say you work at a big tech company. Some rockstar developer's AI agent from the Trailblazer Team drowns you in five dozen "critical vulnerability" tickets for the component you are responsible for.
Do you: a) spend several hours on each ticket to prove that the vulnerability is a hallucination and the "fix" just adds a redundant check - just to get a bad yearly review for "below-average productivity", "not being a team player", and "failing to adjust to the evolving technological landscape".
Or do you b) glance over it, see that it is harmless, and press "Merge" after two minutes with a "LGTM, keep up the good work!" - and get a good yearly review with a raise due to "great cycle time"?
Right, because all these companies and developers are all so nonchalant about pumping out arbitrary code changes, because they have never learned that even small changes cause huge issues despite having experienced it countless times, sometimes to losses of millions of dollars!
There's also c) use your own agent to take the vulnerability and get back to you with an assessment including whether it managed to devise a working exploit, and you go from there.
Or do you have doubts about the ability of these things to devise working exploits? ;-)
usernomdeguerre · · focus · HN ↗
From his Kernel Recipes 2026 slide on Mythos
```
```GHK called this "10 'real' bugfixes", which to me sounds like there's a wild hype machine around these companies and uncritical parroting of every press release they make that falls apart when you engage the affected real experts.
p-o · · focus · HN ↗
OtherShrezzing · · focus · HN ↗
Most of the real low hanging fruit was picked up by humans years ago. When doing automated scanning, the majority of stuff is overly-verbose nonsense which takes hours of expert human labour to understand, test, and discard.
Reading through a Claude generated false positive is absolutely excruciating, because it is absolutely determined that what it’s found is justified. Often you’ll receive very long accompanying “proof of concept” code which demonstrates absolutely wild scenarios. It’s especially frustrating when you’re volunteering your time for a project, and a well-meaning contributor submits the report without the technical nous to understand why you’re rejecting it.
b112 · · focus · HN ↗
Very gung ho, full of energy, loads of book learning, no real world experience or understanding of why things are as they are.
Leave them to their own devices at your peril. Trust nothing they do.
Yet directly guide them, monitor everything they do, some value emerges.
charcircuit · · focus · HN ↗
fc417fc802 · · focus · HN ↗
TeMPOraL · · focus · HN ↗
In human terms, that's already at least a standard deviation above average person.
12376 · · focus · HN ↗
Gigachad · · focus · HN ↗
It’s still good that some real bugs are being patched but what is being reported to the media is so overblown.
crote · · focus · HN ↗
Anthropic claimed that Mythos was so good at finding vulnerabilities that it was too dangerous to release to the public. As this post clearly shows: that (only again) simply isn't true. If you believe your favorite flavor of frontier model is the exception, it is up to you to provide proof to back up that claim.
verdverm · · focus · HN ↗
spwa4 · · focus · HN ↗
1) nuanced tools that talk back, question assumptions, take over decisions, ... oh and expose just how much management knows about the business. Or how little)
2) a tool that can provide the excuse "we've had our source checked and dealt with the remarks"
We all know the answer.
truncate · · focus · HN ↗
b112 · · focus · HN ↗
t_mahmood · · focus · HN ↗
darkwater · · focus · HN ↗
skinfaxi · · focus · HN ↗
delusional · · focus · HN ↗
That's simply not true. I've had some very talented interns, and they are leagues ahead of what the LLMs can do. Not that it matters though, because the point of having interns wasn't to have them produce value. What made the investment worth it was that 12 months down the line I would have a competent colleague that I could have an interesting conversation with. A human person that could challenge some of my blind spots. A person that could take responsibility of something. Maybe not my most important work, but some of it. You don't get ANY of that from the LLM.
Zigurd · · focus · HN ↗
The other side of the coin is that coding agents are not maximally productive unless you give them enough rope to potentially hang themselves. Over roughly the past year, the coding agents I use have gone from hot garbage to pretty consistently useful, especially if I find tasks where I can give them a lot of running room. On the other hand, last week I found a case where the coding agent was looping and flailing like it was doing every third try a year ago.
They fail less often, but they fail in the same way.
bitwize · · focus · HN ↗
There is a vulnerability in library X when you call Y with these specially crafted parameters; as seen in the attached logs it can overflow buffer Z and clobber memory potentially leading to an RCE.
After:
Honest take: the load-bearing constraint violation is real. The log documenting the exploit gate is the ledger which weaves the story.
prox · · focus · HN ↗
bitwize · · focus · HN ↗
IshKebab · · focus · HN ↗
Juliate · · focus · HN ↗
gregw2 · · focus · HN ↗
That said, I remember trying to weigh the hype at the time of the announcement reading/skimming the papers Anthropic published, recognizing that bugcount alone wasn't super-relevant but also remember being impressed by an NFS bug and a kernel bug that struck me as relevant at the time. So where did that NFS issue show up in GKH's list you showed so nicely above?
It turns out, AFAICT, it's not on his list, but the reasons are perhaps interesting to others so I will post here. It turns out there were two NFS issues this past year conflated a bit in my memory:
* The Linux CVE-2026-31402 NFS heap overflow that could allow unauthenticated memory reads over the network isn't in that list of 79, presumably because it was found by Claude Code, not Mythos months earlier. (I am guessing it's not his "malicious network packet into the middle of the stack" and is a stronger attack being a remote attack.)
* And the CVE-2026-4747 NFS stack buffer overflow that allowed gaining full unauthenticated remote root access didn't show up in GKH's list of 79 because despite being Mythos-caught, it wasn't Linux, it was FreeBSD.
I guess this does match my memory now that I think about it, that there weren't any smoking Linux guns caught by Mythos.
* (I guess there was also a longstanding 27-year old OpenBSD TCP SACK-handling stack integer overflow than enabled remote crashes / Denial of Service found by Mythos.)
There is definitely Mythos hype, but just because it hit the BSD code base more than the GKH-managed Linux code base doesn't mean it was inappropriate to raise eyebrows from Mythos, in particular since "attacks only get better".
cyanydeez · · focus · HN ↗
verdverm · · focus · HN ↗
(my current favorite definition)
bigfishrunning · · focus · HN ↗
verdverm · · focus · HN ↗
<a href="https://www.npr.org/transcripts/g-s1-14793" rel="nofollow">https://www.npr.org/transcripts/g-s1-14793
(I think that's the Sherry Turkle episode I'm looking for)
the other I typically reference: <a href="https://www.npr.org/2025/07/18/g-s1177-78041/what-to-do-when-ai-says-i-love-you-we-talked-to-an-artificial-intimacy-expert" rel="nofollow">https://www.npr.org/2025/07/18/g-s1177-78041/what-to-do-when...
darkwater · · focus · HN ↗
Betelbuddy · · focus · HN ↗
malkia · · focus · HN ↗
Betelbuddy · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
FLeXMurphy · · focus · HN ↗
Zip-zapping the bouzouki...
Exfiltrating nuclear arm codes...
Thought for 76 seconds.
You're right to push back on that. That's on me.
catdog · · focus · HN ↗
Mythos turned out to be exactly the marketing stunt it smelled like.
There are others like AISLE who seem to be a bit more successful in finding actual issues using LLMs in some shape or form though, whatever they do differently. Chances are high the secret sauce is not so much about the model being exceptionally powerful which would be bad news for the frontier labs.
duttish · · focus · HN ↗
My understanding: Many many small models in a custom system rather than the biggest and latest
goobreee · · focus · HN ↗
internet_points · · focus · HN ↗
<a href="https://files.mastodon.social/cache/media_attachments/files/117/371/826/895/212/456/original/96ce1f4bfe103ee4.jpeg" rel="nofollow">https://files.mastodon.social/cache/media_attachments/files/...
Faaak · · focus · HN ↗
wslh · · focus · HN ↗
> Any project that has not scanned their source code with AI powered tooling will likely find huge number of flaws, bugs and possible vulnerabilities with this new generation of tools. Mythos will, and so will many of the others.
Greg's video is a good reality check on the hype. But I'd be careful about generalizing from Linux, libcurl, etc which get far more scrutiny than software projects in general. LLM-assisted bug finding still matter a lot for everyday custom and less popular software.
saidnooneever · · focus · HN ↗
mrob · · focus · HN ↗
If you ever used a USB storage device you're vulnerable to this one. Not even a strict chain of custody guarantees safety, because USB devices are often powered by exploitable programmable microcontrollers. If a known good USB device can be converted to a malicious USB device by unprivileged software, the malicious filesystem exploit becomes a local privilege escalation. It works better than tampering with the files on the filesystem because it escapes signature checks and gets you directly into kernel mode.
yubblegum · · focus · HN ↗
TacticalCoder · · focus · HN ↗
It depends on how your host is configured: just as you can do GPU-passthrough, you can passthrough a single USB device or passthrough an entire USB controller to a VM.
[deleted] · · focus · HN ↗
[deleted]
kylestanfield · · focus · HN ↗
sick_of_slop · · focus · HN ↗
[dead]
hodgesrm · · focus · HN ↗
In our company we found real security issues in proprietary code using Opus 4.6/4.7. Obviously typical attackers might have difficulty finding these without code access but Claude was finding real CVEs, which we fixed.
p.s., This is an argument for not trying to deal with security problems by neutering the LLMs. To the extent LLMs are effective it weakens security.
Edit: added p.s.
keeda · · focus · HN ↗
crote · · focus · HN ↗
Let's say you work at a big tech company. Some rockstar developer's AI agent from the Trailblazer Team drowns you in five dozen "critical vulnerability" tickets for the component you are responsible for.
Do you: a) spend several hours on each ticket to prove that the vulnerability is a hallucination and the "fix" just adds a redundant check - just to get a bad yearly review for "below-average productivity", "not being a team player", and "failing to adjust to the evolving technological landscape".
Or do you b) glance over it, see that it is harmless, and press "Merge" after two minutes with a "LGTM, keep up the good work!" - and get a good yearly review with a raise due to "great cycle time"?
keeda · · focus · HN ↗
There's also c) use your own agent to take the vulnerability and get back to you with an assessment including whether it managed to devise a working exploit, and you go from there.
Or do you have doubts about the ability of these things to devise working exploits? ;-)