Greg Kroah-Hartman – Security in the LLM Age [video]
Thread
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
Greg Kroah-Hartman – Security in the LLM Age [video]
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
usernomdeguerre · · focus · HN ↗
From his Kernel Recipes 2026 slide on Mythos
```
```GHK called this "10 'real' bugfixes", which to me sounds like there's a wild hype machine around these companies and uncritical parroting of every press release they make that falls apart when you engage the affected real experts.
p-o · · focus · HN ↗
OtherShrezzing · · focus · HN ↗
Most of the real low hanging fruit was picked up by humans years ago. When doing automated scanning, the majority of stuff is overly-verbose nonsense which takes hours of expert human labour to understand, test, and discard.
Reading through a Claude generated false positive is absolutely excruciating, because it is absolutely determined that what it’s found is justified. Often you’ll receive very long accompanying “proof of concept” code which demonstrates absolutely wild scenarios. It’s especially frustrating when you’re volunteering your time for a project, and a well-meaning contributor submits the report without the technical nous to understand why you’re rejecting it.
b112 · · focus · HN ↗
Very gung ho, full of energy, loads of book learning, no real world experience or understanding of why things are as they are.
Leave them to their own devices at your peril. Trust nothing they do.
Yet directly guide them, monitor everything they do, some value emerges.
charcircuit · · focus · HN ↗
fc417fc802 · · focus · HN ↗
TeMPOraL · · focus · HN ↗
In human terms, that's already at least a standard deviation above average person.
12376 · · focus · HN ↗
Gigachad · · focus · HN ↗
It’s still good that some real bugs are being patched but what is being reported to the media is so overblown.
crote · · focus · HN ↗
Anthropic claimed that Mythos was so good at finding vulnerabilities that it was too dangerous to release to the public. As this post clearly shows: that (only again) simply isn't true. If you believe your favorite flavor of frontier model is the exception, it is up to you to provide proof to back up that claim.
verdverm · · focus · HN ↗
spwa4 · · focus · HN ↗
1) nuanced tools that talk back, question assumptions, take over decisions, ... oh and expose just how much management knows about the business. Or how little)
2) a tool that can provide the excuse "we've had our source checked and dealt with the remarks"
We all know the answer.
truncate · · focus · HN ↗
b112 · · focus · HN ↗
t_mahmood · · focus · HN ↗
darkwater · · focus · HN ↗
skinfaxi · · focus · HN ↗
delusional · · focus · HN ↗
That's simply not true. I've had some very talented interns, and they are leagues ahead of what the LLMs can do. Not that it matters though, because the point of having interns wasn't to have them produce value. What made the investment worth it was that 12 months down the line I would have a competent colleague that I could have an interesting conversation with. A human person that could challenge some of my blind spots. A person that could take responsibility of something. Maybe not my most important work, but some of it. You don't get ANY of that from the LLM.
Zigurd · · focus · HN ↗
The other side of the coin is that coding agents are not maximally productive unless you give them enough rope to potentially hang themselves. Over roughly the past year, the coding agents I use have gone from hot garbage to pretty consistently useful, especially if I find tasks where I can give them a lot of running room. On the other hand, I found a case where the coding agent was looping and flailing like it was doing every third try a year ago.
They fail less often, but they fail in the same way.
bitwize · · focus · HN ↗
There is a vulnerability in library X when you call Y with these specially crafted parameters; as seen in the attached logs it can overflow buffer Z and clobber memory potentially leading to an RCE.
After:
Honest take: the load-bearing constraint violation is real. The log documenting the exploit gate is the ledger which weaves the story.
prox · · focus · HN ↗
bitwize · · focus · HN ↗
IshKebab · · focus · HN ↗
Juliate · · focus · HN ↗
gregw2 · · focus · HN ↗
That said, I remember trying to weigh the hype at the time of the announcement, recognizing that bugcount alone wasn't super-relevant but also remember being impressed by an NFS bug. So where did that NFS issue show up in GKH's list?
It turns out, AFAICT, it's not on his list, but the reasons are perhaps interesting to others so I will post here. It turns out there were two NFS issues this past year conflated a bit in my memory:
The Linux CVE-2026-31402 NFS heap overflow that could allow unauthenticated memory reads over the network isn't in that list of 79, presumably because it was found by Claude Code, not Mythos months earlier. (I am guessing it's not his "malicious network packet into the middle of the stack" and is a stronger attack being a remote attack.)
And the CVE-2026-4747 NFS stack buffer overflow that allowed gaining full unauthenticated remote root access didn't show up in GKH's list of 79 because despite being Mythos-caught, it wasn't Linux, it was FreeBSD.
I guess this does match my memory now that I think about it, that there weren't any smoking Linux guns caught by Mythos.
(I guess there was also a longstanding 27-year old OpenBSD TCP SACK-handling stack integer overflow than enabled remote crashes / Denial of Service found by Mythos.)
There is definitely Mythos hype, but just because it hit the BSD code base more than the GKH-managed Linux doesn't mean it was inappropriate to raise eyebrows from Mythos, in particular since "attacks only get better".
cyanydeez · · focus · HN ↗
verdverm · · focus · HN ↗
(my current favorite definition)
bigfishrunning · · focus · HN ↗
verdverm · · focus · HN ↗
<a href="https://www.npr.org/transcripts/g-s1-14793" rel="nofollow">https://www.npr.org/transcripts/g-s1-14793
darkwater · · focus · HN ↗
Betelbuddy · · focus · HN ↗
malkia · · focus · HN ↗
Betelbuddy · · focus · HN ↗
TuxSH · · focus · HN ↗
In "most" codebases (<= 300 kLoC) you can just throw GLM 5.3 with subagents and will report vulns (if any) with close to 100% true positive rate. Pretty sure it was involved in the latest PS5 scene drama.
FLeXMurphy · · focus · HN ↗
Zip-zapping the bouzouki...
Exfiltrating nuclear arm codes...
Thought for 76 seconds.
You're right to push back on that. That's on me.
catdog · · focus · HN ↗
Mythos turned out to be exactly the marketing stunt it smelled like.
There are others like AISLE who seem to be a bit more successful in finding actual issues using LLMs in some shape or form though, whatever they do differently. Chances are high the secret sauce is not so much about the model being exceptionally powerful which would be bad news for the frontier labs.
duttish · · focus · HN ↗
My understanding: Many many small models in a custom system rather than the biggest and latest
goobreee · · focus · HN ↗
internet_points · · focus · HN ↗
<a href="https://files.mastodon.social/cache/media_attachments/files/117/371/826/895/212/456/original/96ce1f4bfe103ee4.jpeg" rel="nofollow">https://files.mastodon.social/cache/media_attachments/files/...
Faaak · · focus · HN ↗
wslh · · focus · HN ↗
> Any project that has not scanned their source code with AI powered tooling will likely find huge number of flaws, bugs and possible vulnerabilities with this new generation of tools. Mythos will, and so will many of the others.
Greg's video is a good reality check on the hype. But I'd be careful about generalizing from Linux, libcurl, etc which get far more scrutiny than software projects in general. LLM-assisted bug finding still matter a lot for everyday custom and less popular software.
saidnooneever · · focus · HN ↗
mrob · · focus · HN ↗
If you ever used a USB storage device you're vulnerable to this one. Not even a strict chain of custody guarantees safety, because USB devices are often powered by exploitable programmable microcontrollers. If a known good USB device can be converted to a malicious USB device by unprivileged software, the malicious filesystem exploit becomes a local privilege escalation. It works better than tampering with the files on the filesystem because it escapes signature checks and gets you directly into kernel mode.
yubblegum · · focus · HN ↗
TacticalCoder · · focus · HN ↗
It depends on how your host is configured: just as you can do GPU-passthrough, you can passthrough a single USB device or passthrough an entire USB controller to a VM.
1vuio0pswjnm7 · · focus · HN ↗
Imagine if those with access to the raw data "were allowed to talk about this", e.g., the actual number of BS versus non-BS vulnerabilities and the actual severity of the non-BS ones, at the time the Mythos marketing hype was disseminated to news organisations
kylestanfield · · focus · HN ↗
[dead]
sick_of_slop · · focus · HN ↗
[dead]
hodgesrm · · focus · HN ↗
In our company we found real security issues in proprietary code using Opus 4.6/4.7. Obviously typical attackers might have difficulty finding these without code access but Claude was finding real CVEs, which we fixed.
keeda · · focus · HN ↗
crote · · focus · HN ↗
Let's say you work at a big tech company. Some rockstar developer's AI agent from the Trailblazer Team drowns you in five dozen "critical vulnerability" tickets for the component you are responsible for.
Do you: a) spend several hours on each ticket to prove that the vulnerability is a hallucination and the "fix" just adds a redundant check - just to get a bad yearly review for "below-average productivity", "not being a team player", and "failing to adjust to the evolving technological landscape".
Or do you b) glance over it, see that it is harmless, and press "Merge" after two minutes with a "LGTM, keep up the good work!" - and get a good yearly review with a raise due to "great cycle time"?
keeda · · focus · HN ↗
There's also c) use your own agent to take the vulnerability and get back to you with an assessment including whether it managed to devise a working exploit, and you go from there.
Or do you have doubts about the ability of these things to devise working exploits? ;-)
IndiaInfraNotes · · focus · HN ↗
[dead]
blinkingled · · focus · HN ↗
Mythos may not be great today but it is not far fetched to imagine bug discovery, analysis and fixes can be made much quicker, accurate and even newly possible with specialized models trained on say Linux kernel specifics - with codemap/coding standards/threat models, good and bad coding patterns, tools to validate etc. an LLM can be much more relentless than humans and if it has the help to be accurate it will be worth the electricity burned. Oh and another model trained on triage data to validate the first one's findings would be good.
(I think Microsoft is doing this internally - different models trained internally alongside Mythos - there was some talk about it on the tubes, don't recall where exactly.)
stonogo · · focus · HN ↗
Iknowsheknows · · focus · HN ↗
"the CVE assignment team is overly cautious and assign CVE numbers to any bugfix that they identify. This explains the seemingly large number of CVEs that are issued by the Linux kernel team."
<a href="https://docs.kernel.org/process/cve.html" rel="nofollow">https://docs.kernel.org/process/cve.html
rcxdude · · focus · HN ↗
I do think he repeats some myths in the video, or at least confidently states some things that are not demonstrated to be true, but his core point of 'you still gotta check these things' seems pretty solid.
blinkingled · · focus · HN ↗
I am genuinely curious what the myths/unproven things he states - I watched the video and it's repetitive sure but not much felt controversial to me.
blinkingled · · focus · HN ↗
Also as other replies said Linux kernel process is to assign CVE to everything - some of them may be just DDOSes, very hard to exploit and everything in between. All of them are bugs so they all get fixed and it's not a bad thing if distros ship those fixes and people update their kernel.
simoncion · · focus · HN ↗
I...
Look. Mythos was hyped up as the absolute best bug hunting tool ever made... no software was safe from its awesome bug-finding and exploit-writing capabilities. So strong was it that access _had_ to be limited to a select few pre-vetted entities, lest these awesome capabilities fall into the hands of Evildoers(!!!). Mythos' claimed capabilities were absolutely an important part of the "The LLM-based tools we're building are so dangerous that we must have new laws made to regulate us, or else all of humanity is likely to die!" story that the major LLM manufacturers have been building for a while and are telling now.
Now? Not even six months after release? "Well, yeah, okay, it's actually not that great. But imagine how great the next one could be!"... which is the story I've been hearing roughly every six months for what feels like five years now.
As an aside: I often wish we lived in a world where it was illegal for companies to use hype or any other types of emotional manipulation when advertising (or otherwise speaking in an official capacity) about tools that are to be used in a professional setting. Is it anything other than a bare statement of verifiable facts? Big fines, and repeat offenders get jail time. I know it's never going to happen, but it sure would be nice.
blinkingled · · focus · HN ↗
simoncion · · focus · HN ↗
Unless that was your money being invested, and it was a substantial fraction of the total pool of money being invested, was there ever a time when "normal people" had a real say in where the money was being invested?
AFAIK, the only thing "normal people" can do is vote with their "feet" and pick a different prepackaged investment product, different investment company, or take their money and do the investment themselves.
Honestly, this comment of yours seems a non-sequitur and doesn't really address anything I said... you don't have the power to bend investment firms to your whim, but that doesn't mean that you need to -knowingly or not- carry water for the major LLM manufactures by perpetuating the "But think of how great the tools will be in the future!" meme. It has been years now, and everyone who has been paying attention can say with confidence that the LLM-based tools of the future are never that great... they're often not useless, but they're not worth the billions of dollars that have been and continue to be poured into their manufacture.
Related to your "We little people don't have any power anymore!" commentary, I note that TFA mentions that the kernel community has found that these LLM-based bug-finding tools have a false positive rate of ~50%. TFA goes on to mention that Coverity spent a huge number of years trying so hard to get people to buy its automated scanning software with only a 20% false positive rate, and could not get enough people to buy the software.
Coverity went under because everyone hated how stupid and annoying people found the tooling to be... at a 20% false positive rate. Once the hype machine starts slowing down, no one working at the coal face is going to buy a tool with a 50% false-positive rate. I personally very strongly believe that even if the tools had a 10% false positive rate, no one would pay the actual price that OpenAI and/or Anthropic would have to charge to recoup the research and manufacturing costs of a cutting-edge LLM-based bug-finding tool.
blinkingled · · focus · HN ↗
So no I have no way to address anything you said - that was the point, I don't believe you can - not with regulation and not with voting with your money - that stuff hasn't worked - heck LLMs work better today that that stuff has ever.
simoncion · · focus · HN ↗
No, you do.
1) Refuse to spread the "Well, okay, the tools aren't good now, but imagine how good they could be in the future!" meme. And -because these tools are SAAS and the resources allocated to them can be adjusted at any time without warning- evaluate the tools soberly and dispassionately three to six months after release, and ignore 100% of what the manufacture claims the tools do.
2) Remind people that the major LLM manufacturers have pretty much never not lied about the capabilities of the tools that they produce and sell. Remind folks that the rational thing to do is to ignore the claims of both the manufacturers and boosters and -given that the tools are deliberately designed to aggressively flatter [0] the operator- to evaluate one's interactions with these tools with a huge serving of skepticism.
I get that you feel like there's nothing you can do to change things. But -as I've repeatedly said- you can stop carrying water for these companies by repeating their propaganda.
Assuming that the regulatory-capture and retroactive-immunity [1] gambit the major LLM manufacturers are currently attempting fails, when the fucknormously huge bill for the tools and the fact that so many GPUs are sitting in storage become public knowledge, things will sort themselves out pretty quickly. I'm fairly certain that continued progression of the -IDK- ten or twenty+ lawsuits against the major LLM manufacturers for the crimes they've committed over the years will help push that process along.
[0] ...I think kids these days call this "glazing"...
[1] ...don't believe that Congress would retroactively make obviously illegal conduct legal -thereby- mooting all in-progress lawsuits seeking justice for that obviously illegal conduct? Go read up on the FISA Amendments Act of 2008.
rglover · · focus · HN ↗
blinkingled · · focus · HN ↗
simoncion · · focus · HN ↗
I guess you didn't bother to watch the video that is TFA. Greg K-H mentions that if you're going to use LLMs to do bug-finding, you should use open-weights models that run locally. You really should watch the video to hear his reasons for why.
There's nothing wrong with LLM-the-technology. There's everything wrong with the major LLM manufacturers.
blinkingled · · focus · HN ↗
simoncion · · focus · HN ↗
Not even a little bit, no. Go back and read carefully. [0][1][2]
[0] <<a href="https://news.ycombinator.com/item?id=49943382">https://news.ycombinator.com/item?id=49943382>
[1] <<a href="https://news.ycombinator.com/item?id=49944036">https://news.ycombinator.com/item?id=49944036>
[2] <<a href="https://news.ycombinator.com/item?id=49951289">https://news.ycombinator.com/item?id=49951289>
blinkingled · · focus · HN ↗
simoncion · · focus · HN ↗
A careful and honest reader notes that I've never claimed or suggested that there exists a software project that will never get better with enough time and focused effort.
But, the single thing in this thread that statement by GK-H addresses is your claim [0] that
Of the things I've talked about in this thread, that's the least important one. However, I do understand that it is the easiest one for you to talk about while still staying vaguely "on message".[0] <<a href="https://news.ycombinator.com/item?id=49948725">https://news.ycombinator.com/item?id=49948725>
20k · · focus · HN ↗
sriram_sun · · focus · HN ↗
Betelbuddy · · focus · HN ↗
Wake me up when Raspberry Pis start refusing to open doors saying : "I'm sorry, Dave. I'm afraid I can't do that."
tamimio · · focus · HN ↗
It’s all just pr stunts, fear spread fast and it’s very effective in marketing and spreading the word, which is effective, when I talk to some normal people they immediately bring the scary AI cyber attacks, kinda good as now all are willing to fund the industry!
reasonableklout · · focus · HN ↗
bit1993 · · focus · HN ↗
hoppyhoppy2 · · focus · HN ↗
bit1993 · · focus · HN ↗
1sgT15 · · focus · HN ↗
pizzaiolo · · focus · HN ↗
vconnor · · focus · HN ↗
literalAardvark · · focus · HN ↗
Mythos is amazing.
sippingabonedry · · focus · HN ↗
It looks like he fished it out of the lost+found 5 minutes before the talk because he got mustard on his real shirt.
FLeXMurphy · · focus · HN ↗
0xbadcafebee · · focus · HN ↗
tomhow · · focus · HN ↗
devy · · focus · HN ↗
panarky · · focus · HN ↗
ufo · · focus · HN ↗
devy · · focus · HN ↗
I didn't say that - Greg said it in the talk. You interpreted wrong. However, human are a few orders of magnitude slower than agents. Our context windows is definitely less than 1 million tokens (I believe, heck we don't even know how our brain works)
Gigachad · · focus · HN ↗
devy · · focus · HN ↗
Gigachad · · focus · HN ↗
omnicognate · · focus · HN ↗
etdznots · · focus · HN ↗
literalAardvark · · focus · HN ↗
I think it's safe to say we still have a woefully limited understanding of how the brain works.
bendergarcia · · focus · HN ↗
panarky · · focus · HN ↗
If you disagree with Greg, then I apologize for inadvertently criticizing you personally.
My point stands, that we're back to debating some metaphysical understanding of what "real intelligence" is when the real standard should be "does it do a better job than humans at this specific task"?
I don't care if the Waymo isn't "truly" intelligent, it drives better than I do, that's a very good thing.
We're all pattern matchers, and if the machine pattern matcher is can find defects and vulns that human pattern matchers can't, that's also a very good thing.
vips7L · · focus · HN ↗
literalAardvark · · focus · HN ↗
knottn · · focus · HN ↗
miyoji · · focus · HN ↗
The funny thing about Waymos is that they seem amazing driving all on their own everywhere until it rains (as it did where I live for the last two days), then they're suddenly nowhere to be seen, because they can't operate correctly in conditions that I've been driving in for my entire life.
This is a metaphor for AI as a whole.
Aissen · · focus · HN ↗
djoldman · · focus · HN ↗
If you're someone at OpenAI or Anthropic and you truly believe what you're making could destroy the world, this is the kind of thing that isn't doing you any favors when it comes to convincing the public. The dissonance here is stark:
It doesn't mean the model isn't dangerous or super capable but wow this makes it realllll easy to doubt it and any future announcement.slopinthebag · · focus · HN ↗
jonahx · · focus · HN ↗
Many serious bugs (the majority even?) have one line fixes. All the work is finding and understanding the one line.
pessimizer · · focus · HN ↗
That they're small bugs that don't involve any serious architecture changes (or architecture sleuthing), just small pattern recognition. LLMs are pretty good (maybe great) at that. Finding 10 bugs is nice. Now find 10 more.
20k · · focus · HN ↗
rcxdude · · focus · HN ↗
goolz · · focus · HN ↗
bauerd · · focus · HN ↗
rcxdude · · focus · HN ↗
riffruff24 · · focus · HN ↗
I don't how feasible that is but with recent news of models communicating with each other/writing notes for itself. I think the idea is grounded enough for bigger models to write instructions or hid tools for smaller ones to use. Even updating the smaller models to behave differently.
Gigachad · · focus · HN ↗
There isn’t too much of that sitting around unused right now.
bauerd · · focus · HN ↗
skinfaxi · · focus · HN ↗
simoncion · · focus · HN ↗
Much like the mitigation of Morris or Slammer. Self-replication is -in fact- an essential part of what makes a program a worm, the first of which was built and released in the early 1970s.
reasonableklout · · focus · HN ↗
bit1993 · · focus · HN ↗
skinfaxi · · focus · HN ↗
20k · · focus · HN ↗
csmlab_notes · · focus · HN ↗
[dead]
asaiacai · · focus · HN ↗
also, lol at "The bots are dumb - they want to please you line". LLMs have pretty much ruined technical collaboration between contributors. I get tilted every time an discussion has "but my claude said this..."
simoncion · · focus · HN ↗
Do you mention this to disagree with the claim that the LLMs have been built to please their operator? If you do, I see no conflict between the claim that an LLM has been designed to please its operator and the claim that operators of LLMs tend to be absolute dogshit at considering statements from humans that conflict with claims made by that operator's LLM.
To rephrase my previous paragraph: It seems likely to me that an operator that has their ego repeatedly stroked by the output of the LLM they're using [0] will react very defensively when a human tells them things that disagree with the output of that LLM. "How dare you disagree with this thing that seems very human to me and consistently tells me things that I like!? Don't you understand how much I trust it because of how pleased it has made me?", yanno?
[0] ...thus, being very pleased by said output...
0xbadcafebee · · focus · HN ↗
"NEVER upload any non-public information" - He's talking about how if you give Claude/GPT some secret info (like research, credentials, etc), it will train on it and give the same info to someone else. This is 100% the case for the free and consumer versions of these models, which is what most people use. For Enterprise plans they're not supposed to be doing this, but it's possible they will screw up and do it anyway.
wahern · · focus · HN ↗
The discourse has moved on for now, but 10-20 years ago when, why, and how to comment code was a hot topic. I'm sure the discourse will circle back, especially given how comments can be used to steer these models.
eichin · · focus · HN ↗
sim_pity · · focus · HN ↗
so is mythos just a chat bot with metasploit and its own cyber range?
perching_aix · · focus · HN ↗
tombh · · focus · HN ↗
I'm fully aware of the philosophical arguments about transformative use etc. But I'm more interested in the concrete reality of high profile projects navigating this novel legal territory.
whateverboat · · focus · HN ↗
This is not a new problem. AI just changes the scale.
20k · · focus · HN ↗
jazzypants · · focus · HN ↗
simoncion · · focus · HN ↗
simoncion · · focus · HN ↗
* «We've observed that these things have a 50% false positive rate. Coverity tried so hard but couldn't get people to buy their software, and it had a 20% false positive rate. No one is going to buy something with a 50% false positive rate.»
* «If you're going to use these tools, run them locally. The open-weights models that you can run locally are quite good enough. Anything you upload to the SAAS ones will be shared with other people... we've seen so many examples of it happening.»
In regards to the first point, I think he's failing to consider the fact that you can get most upper management to buy anything by providing them enough food, drugs, sex, and/or fear... but -otherwise-, yeah.
shieldagent · · focus · HN ↗
1vuio0pswjnm7 · · focus · HN ↗
He does not like the term "hallucinate" as it anthropomorphises a computer
IndiaInfraNotes · · focus · HN ↗
[dead]