Greatly appreciated the candor. I've included a few slides into text that i thought were eye-opening to me:
From his Kernel Recipes 2026 slide on Mythos
```
Mythos's 79 vulnerabilities:
24 - no detail at all "something crashed"
14 - not a bug at all
3 - totally made up data
15 - already fixed in latest release
- 11 by others
- 4 by anthropic
20 - fixes were needed
- 7 "assume a malicious filesystem image"
- 2 "assume you can inject a malicious network packet into the middle of the stack"
- 2 "NOMMU"
- 6 sctp networking issues for untrusted devices
- 2 ipv6 minor network issues
- 1 gpu driver for local malicious user
```
GHK called this "10 'real' bugfixes", which to me sounds like there's a wild hype machine around these companies and uncritical parroting of every press release they make that falls apart when you engage the affected real experts.
We’ve seen this in a few open source repos we voluntarily manage security on. They’re not massive repos, but big enough they get attention from security researchers.
Most of the real low hanging fruit was picked up by humans years ago. When doing automated scanning, the majority of stuff is overly-verbose nonsense which takes hours of expert human labour to understand, test, and discard.
Reading through a Claude generated false positive is absolutely excruciating, because it is absolutely determined that what it’s found is justified. Often you’ll receive very long accompanying “proof of concept” code which demonstrates absolutely wild scenarios. It’s especially frustrating when you’re volunteering your time for a project, and a well-meaning contributor submits the report without the technical nous to understand why you’re rejecting it.
In some ways the reality is worse: the same version of an LLM won't get better at its job, even though you might get better at prompting it. Newer versions are trained on more code, which has obvious benefits, and their harnesses are better at taking advantage of tools that were created to keep human coders out of trouble.
The other side of the coin is that coding agents are not maximally productive unless you give them enough rope to potentially hang themselves. Over roughly the past year, the coding agents I use have gone from hot garbage to pretty consistently useful, especially if I find tasks where I can give them a lot of running room. On the other hand, last week I found a case where the coding agent was looping and flailing like it was doing every third try a year ago.
They fail less often, but they fail in the same way.
usernomdeguerre · · focus · HN ↗
From his Kernel Recipes 2026 slide on Mythos
```
```GHK called this "10 'real' bugfixes", which to me sounds like there's a wild hype machine around these companies and uncritical parroting of every press release they make that falls apart when you engage the affected real experts.
OtherShrezzing · · focus · HN ↗
Most of the real low hanging fruit was picked up by humans years ago. When doing automated scanning, the majority of stuff is overly-verbose nonsense which takes hours of expert human labour to understand, test, and discard.
Reading through a Claude generated false positive is absolutely excruciating, because it is absolutely determined that what it’s found is justified. Often you’ll receive very long accompanying “proof of concept” code which demonstrates absolutely wild scenarios. It’s especially frustrating when you’re volunteering your time for a project, and a well-meaning contributor submits the report without the technical nous to understand why you’re rejecting it.
b112 · · focus · HN ↗
Very gung ho, full of energy, loads of book learning, no real world experience or understanding of why things are as they are.
Leave them to their own devices at your peril. Trust nothing they do.
Yet directly guide them, monitor everything they do, some value emerges.
Zigurd · · focus · HN ↗
The other side of the coin is that coding agents are not maximally productive unless you give them enough rope to potentially hang themselves. Over roughly the past year, the coding agents I use have gone from hot garbage to pretty consistently useful, especially if I find tasks where I can give them a lot of running room. On the other hand, last week I found a case where the coding agent was looping and flailing like it was doing every third try a year ago.
They fail less often, but they fail in the same way.