Several vulnerabilities have been discovered in the Linux kernel
Thread
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
Several vulnerabilities have been discovered in the Linux kernel
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
modeless · · focus · HN ↗
nathell · · focus · HN ↗
I suggest this post be renamed “A legion of vulnerabilities has been discovered…”
rotis · · focus · HN ↗
pluc · · focus · HN ↗
mrmuagi · · focus · HN ↗
kachnuv_ocasek · · focus · HN ↗
For context, there are 48,152 CVEs from 2025 and 73,356 in 2026 so far.
devy · · focus · HN ↗
Yeah, the "Several" is a massive understatement.
thallium205 · · focus · HN ↗
wjholden · · focus · HN ↗
vdfs · · focus · HN ↗
slopinthebag · · focus · HN ↗
one again illustrating the importance of encapsulating unsafe behavior. perhaps c should get a __UNSAFE { } block, where memory access is encapsulated and thus most bugs occurring outside of those blocks do not need to be marked as CVEs.
akersten · · focus · HN ↗
I think the convention for this is at the filesystem level and most programmers use the `.c` suffix to indicate it
slopinthebag · · focus · HN ↗
perhaps we could call it a sedimentchest?
catlifeonmars · · focus · HN ↗
someonebaggy · · focus · HN ↗
branc116 · · focus · HN ↗
insanitybit · · focus · HN ↗
debugnik · · focus · HN ↗
seba_dos1 · · focus · HN ↗
ofjcihen · · focus · HN ↗
That said, 1,313 CVEs in one advisory probably says as much about how aggressively the kernel tracks and batches fixes as it does about Linux suddenly becoming uniquely insecure. Still a pretty incredible number to see in one place.
DominoTree · · focus · HN ↗
ofjcihen · · focus · HN ↗
[dead]
BobbyTables2 · · focus · HN ↗
Seems like an enormous increase over 2024 and 2025.
ganelonhb · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
tetrisgm · · focus · HN ↗
SchemaLoad · · focus · HN ↗
catlifeonmars · · focus · HN ↗
skeledrew · · focus · HN ↗
Wouldn't this mean people are actually encountering issues where there are none before? Likely some serious enough that they lead to exploits where there weren't before? Where are the reports of these new defects?
catlifeonmars · · focus · HN ↗
somenameforme · · focus · HN ↗
Long term I suspect that the purpose of the digital domain is going to end up being rethought. For instance connecting critical infrastructure to the internet has always been a terrible idea, and LLMs will just make that even more clear.
SchemaLoad · · focus · HN ↗
On the iphone for example even if you find a crippling bug in iOS which gives you full root access, there is still no way to get the device encryption key or face ID info because the secure enclave simply has no electrical connection that can pass that key to the OS.
kamma4434 · · focus · HN ↗
And this is the best of cases. I fear many times in a smaller projects it was “it compiles at last, let’s see if anybody complains”. That’s why LLM’s are so damn effective today.
(been there, done that – I’m not pointing fingers, but we are human, we get tired, and we don’t have the NASA budget to complete things and ship them)
sva_ · · focus · HN ↗
Apparently by late summer there were already more vulnerabilities found this year than in all of 2025.
cryptonym · · focus · HN ↗
SadErn · · focus · HN ↗
oridjeicjejdj · · focus · HN ↗
rockskon · · focus · HN ↗
They're not demanded because targets are just so easy to pop.
If we did secure software across the board with AI, there'd likely be a resurgence of calls for mandatory backdoors.
autoexec · · focus · HN ↗
imoverclocked · · focus · HN ↗
crtasm · · focus · HN ↗
<a href="https://security-tracker.debian.org/tracker/source-package/linux" rel="nofollow">https://security-tracker.debian.org/tracker/source-package/l...
imoverclocked · · focus · HN ↗
Searching around, the best I have found so far for vanilla kernels is: <a href="https://linuxcvetracker.com" rel="nofollow">https://linuxcvetracker.com
It does require a little clicking around to get all the info I want though. Time to pull out curl+awk! :)
evgpbfhnr · · focus · HN ↗
2OEH8eoCRo0 · · focus · HN ↗
embedding-shape · · focus · HN ↗
SchemaLoad · · focus · HN ↗
Now they just give almost every bug a CVE number.
vdfs · · focus · HN ↗
crispr245 · · focus · HN ↗
tclancy · · focus · HN ↗
rurban · · focus · HN ↗
zahlman · · focus · HN ↗
> In the Linux kernel, the following vulnerability has been resolved: > usb: typec: ucsi: unregister debugfs entries on teardown > ucsi_register() creates per-instance debugfs entries, but > ucsi_unregister() keeps them around until ucsi_destroy(). > Drivers like ucsi_glink that unregister/register the same UCSI > instance across remoteproc restart then try to create an already > existing debugfs directory and log: > debugfs: 'pmic_glink.ucsi.0' already exists in 'ucsi' > Unregister debugfs entries as part of ucsi_unregister(), and > clear ucsi->debugfs after freeing it so repeated unregister > paths remain safe.
I'm going to need someone to explain how that could possibly become a "vulnerability".
brabel · · focus · HN ↗
jaimex2 · · focus · HN ↗
radicalcentrist · · focus · HN ↗
infthi · · focus · HN ↗
userbinator · · focus · HN ↗
Remotely or locally exploitable? This is very lacking on information.
SchemaLoad · · focus · HN ↗
adastra22 · · focus · HN ↗
walrus01 · · focus · HN ↗
john_strinlai · · focus · HN ↗
>“Due to the layer at which the Linux kernel is in a system, almost any bug might be exploitable to compromise the security of the kernel… Because of this, the CVE assignment team is overly cautious and assign CVE numbers to any bugfix that they identify.”
<a href="https://docs.kernel.org/process/cve.html" rel="nofollow">https://docs.kernel.org/process/cve.html
"number of cves" is a useless metric, especially when it comes to the kernel.
rerdavies · · focus · HN ↗
SoftTalker · · focus · HN ↗
catlifeonmars · · focus · HN ↗
SoftTalker · · focus · HN ↗
odo1242 · · focus · HN ↗
rerdavies · · focus · HN ↗
:-P
whiskey-one · · focus · HN ↗
catlifeonmars · · focus · HN ↗
bigstrat2003 · · focus · HN ↗
PaulDavisThe1st · · focus · HN ↗
The user may have no idea that the plugin is malicious; the program remains bug-free (if it was beforehand).
nananana9 · · focus · HN ↗
Much more agregiously you can design harmless looking formats that can run arbitrary code (e.g. .doc with VBS). I have a very hard time blaming an user who falls for that enen though MS puts up a scary looking popup.
GoblinSlayer · · focus · HN ↗
PowerElectronix · · focus · HN ↗
SAI_Peregrinus · · focus · HN ↗
If you're willing to stretch, missing but planned features also deny the use of said features since they haven't been added yet, and so are CVSS 1/Low vulnerabilities.
Resume-driven development for security researchers has never been easier!
viraptor · · focus · HN ↗
That doesn't follow. In the extremely simple example, an adding service returning 1+1=3 has a bug, but it's not a possible DoS situation at all.
> missing but planned features also deny the use of said features
That's not what DoS is.
This whole situation with CVE assigning comes from the whole process being far from ideal. But it doesn't mean it's completely useless and doesn't follow any rules at all.
Gigachad · · focus · HN ↗
Until someone finds there is a user input they can trigger this bug causing some other bit of code to read data from the wrong offset and now it's a whole exploit.
viraptor · · focus · HN ↗
someonebaggy · · focus · HN ↗
tsimionescu · · focus · HN ↗
asdfaoeu · · focus · HN ↗
asdfaoeu · · focus · HN ↗
viraptor · · focus · HN ↗
nikanj · · focus · HN ↗
eastbound · · focus · HN ↗
josephg · · focus · HN ↗
viraptor · · focus · HN ↗
nikanj · · focus · HN ↗
<a href="https://daniel.haxx.se/blog/2026/06/24/a-cve-dispute/" rel="nofollow">https://daniel.haxx.se/blog/2026/06/24/a-cve-dispute/
literalAardvark · · focus · HN ↗
Organizations that track their software and patch systems for CVEs have their own risk management system.
Publish xlow as low, we'll filter them out if we want to. As even they state in that article, they do NOT have the necessary context to filter stuff out. So why are they doing it?
woodruffw · · focus · HN ↗
kbolino · · focus · HN ↗
What does this even mean? cURL is one of the most load-bearing pieces of software in existence. It, and the Linux kernel, which takes a similarly dim view of the CVE process, are the inventors. "Not Invented Here" seems to imply that there is a vast body of peer work for them to draw on to resolve this problem, but who are their peers? As far as I can tell, the answer is "Microsoft and Apple" who exist in a totally different, mostly closed-source or at least closed-development, ecosystem.
dzhiurgis · · focus · HN ↗
Is it really, or it's trivially replaceable by wget?
kbolino · · focus · HN ↗
Regardless, even if hypothetical product X could do everything cURL and libcurl do for most people, that wouldn't make cURL/libcurl any less load-bearing, any more than the existence of OpenBSD or Illumos make Linux less load-bearing.
viraptor · · focus · HN ↗
and one that will never make decisions that is incorrect or disliked by any party.
nikanj · · focus · HN ↗
viraptor · · focus · HN ↗
euank · · focus · HN ↗
Read <a href="http://www.kroah.com/log/blog/2026/02/16/linux-cve-assignment-process/" rel="nofollow">http://www.kroah.com/log/blog/2026/02/16/linux-cve-assignmen...
Since Linux can't really know exactly how it will be used, it's almost impossible to accurately assess the impact of bugs.
Like, maybe 6 million closed-source applications running on linux have:
Which would make that bug a DoS.Linux is offering a general use tool, so they can't know exactly how people use it, so they can't know what will and won't be severe or not in terms of security impact.
It's not worth their time to triage what users of their kernel are doing for every bug and figure out if it's a security issue or not.
goodmythical · · focus · HN ↗
Resulting from miscalc of either stored or supplied would indeed create DoS.
If the off by one is in a graphical driver that causes an overflow leading to no output, that's DoS.
If the off by one is in the memory mapping of input devices leading to no available input, that's DoS.
Simple math is kind of everywhere in the kernel and userland apps. If the math is wrong and results in memory mapping wrong such that kernel panics or is unuseable, that's a breaking bug regardless of simplicity.
viraptor · · focus · HN ↗
I'm using specific words and it's been explained and linked twice already. Come on!
jeroenhd · · focus · HN ↗
However, the Linux kernel is supposed to run any userland program without crashing, so anything that crashes the kernel is a local DoS and there are a lot of them. It's also supposed to shield processes from each other and maintain privilege levels correctly, so many incorrect memory leaks are also CVE worthy. Whether a CVE applies depends on the people and programs using the kernel, and the kernel team can't read your code to tell you if it applies or not.
People reading CVEs wrong ("it's got a high number so we must patch within a day") must be going crazy over this, but the point of CVEs is to let you make judgement calls, not to be a cool statistic about how secure something is.
Most CVEs are irrelevant to most people, that's always been the case.
somat · · focus · HN ↗
Unpatched bug, wrong color lights.
Yes, it is a bit of a stretch, I desperately hope programmable navigation lights are not a thing. And I also don't think every bug needs a CVE. But in the correct context nearly any bug could be critical.
seanhunter · · focus · HN ↗
[1] <a href="https://www.computerweekly.com/news/1280091718/Chinook-computer-was-positively-dangerous-say-newly-disclosed-MoD-documents" rel="nofollow">https://www.computerweekly.com/news/1280091718/Chinook-compu...
rjsw · · focus · HN ↗
There have been documented problems with Windows for Warships.
seanhunter · · focus · HN ↗
sas224dbm · · focus · HN ↗
jeroenhd · · focus · HN ↗
MyMemoryfails · · focus · HN ↗
GTP · · focus · HN ↗
sigmoid10 · · focus · HN ↗
TeMPOraL · · focus · HN ↗
The vendors fixing them arguably prioritize these reports right. Most of the CVEs, even severe ones, are irrelevant in practice, and as parents note, are more like regular bugs with security flavor in reporting. The CVE label instead of regular bug tracking number makes them seem important.
openasocket · · focus · HN ↗
But I think the most important thing to keep in mind is that a bug isn’t necessarily less serious or less important than a vulnerability. A serious bug should be patched just as urgently as a serious vulnerability.
funcDropShadow · · focus · HN ↗
literalAardvark · · focus · HN ↗
twoodfin · · focus · HN ↗
Sesse__ · · focus · HN ↗
Realistically, most admins cannot make judgment calls about 1000+ CVEs for a kernel release.
jurgenburgen · · focus · HN ↗
lstodd · · focus · HN ↗
zero CVE policy = halt on business development.
dormento · · focus · HN ↗
Even more realistically, many admins do not have the background to be able to reason (by themselves) about the actual risk of most CVEs, so just going along with specialized media coverage is often a sound strategy.
wang_li · · focus · HN ↗
We should keep our hopes up, someday we may get there.
spragl · · focus · HN ↗
[dead]
theteapot · · focus · HN ↗
mbreese · · focus · HN ↗
I do find it interesting though, that in the interest of transparency, every bugfix gets a CVE. Which ends up being a huge number… which will ultimately yield a more insecure environment as we’re getting conditioned to ignore/discount CVEs by the volume.
Over-reporting in this case seems to risk being counterproductive.
Gigachad · · focus · HN ↗
The frequency and severity of cyber attacks has increased to the point a much more cautious approach has become common. It's also easier to sell this work to management when you can point at the security tab on some tool and say "Look we need to patch these CVEs"
autoexec · · focus · HN ↗
Gigachad · · focus · HN ↗
1718627440 · · focus · HN ↗
Gigachad · · focus · HN ↗
1718627440 · · focus · HN ↗
seb1204 · · focus · HN ↗
fwip · · focus · HN ↗
izacus · · focus · HN ↗
cesarb · · focus · HN ↗
> [...] a much more cautious approach has become common.
I'd argue that this is a less cautious approach, not more. It takes time to carefully evaluate, review, and test each change.
socializer · · focus · HN ↗
Kernel development is well-funded, both via grants and by direct employment at big tech companies, and if they wanted to properly triage and annotate vulnerabilities, and provide reasonable assessments of what is or isn't likely to be a security risk, they absolutely could. They almost certainly could go to Google and say "we need two people full-time on your payroll for this" and they would get it.
I don't want to dunk on them too much because they're generally doing God's work, but these absolutist security stances are not worth being taken seriously.
It's basically saying that they can't possibly provide a valuable service for 99.999% of the install base because there might a hypothetical person out there using Linux in a really weird way. If Microsoft tried to make an argument like that, they'd get crucified.
vlovich123 · · focus · HN ↗
seb1204 · · focus · HN ↗
cannonpalms · · focus · HN ↗
asdfaoeu · · focus · HN ↗
serbuvlad · · focus · HN ↗
Linux only ever wanted to promise support for the latest release and even Linux LTS is a concession.
And CVEs are basically a useless concept if you live at HEAD. (or at least not any more useful than any other bug tracker which supports tags)
LtWorf · · focus · HN ↗
If he kept true of his "we don't break user space" instead of it being "we don't break user space until we do and then it's on you to deal with it" perhaps more people would be willing to run the latest release.
serbuvlad · · focus · HN ↗
Linux LTS is much more about proprietary drivers targeting a stable internal kernel API/ABI than about anything else.
LtWorf · · focus · HN ↗
mwwaters · · focus · HN ↗
Internal Kernel API changes all the time which can break proprietary drivers. But userspace API has a far higher guarantee.
LtWorf · · focus · HN ↗
doublepg23 · · focus · HN ↗
throwaway7356 · · focus · HN ↗
LtWorf · · focus · HN ↗
DSMan195276 · · focus · HN ↗
I'll challenge this, I don't think that's really possible, at least not to any high degree of confidence. Nobody can realistically evaluate if any particular out-of-bounds read/write or use-after-free is "safe", and if you're going to consider all of those as security risks then there's not really a point in trying to filter out the few bug fixes that might not lead to those things.
Try going to the linked page, pick any random CVE, and read it. I've checked a bunch and I'd say _at least_ 8/10 of them are variations on those two things.
someonebaggy · · focus · HN ↗
weinzierl · · focus · HN ↗
st_goliath · · focus · HN ↗
Any patch that is back ported to a stable kernel, indiscriminately. And they have also started assigning CVSS scores with the same kind of malicious compliance.
Take for example, this patch in the device mapper RAID code:
<a href="https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=47f1441b281decde6954a2fa82b4131637d685ac" rel="nofollow">https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/lin...
After back porting to stable, it got assigned CVE-2026-89558 (which is in this list), and a CVSS score of 9.8:
<a href="https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2026/CVE-2026-89558.cvss" rel="nofollow">https://git.kernel.org/pub/scm/linux/security/vulns.git/tree...
Reasoning behind it being that theoretically, a RAID could be accessible over the network via NFS, iSCSI, etc... so if it gets corrupted, the buggy code path in the recovery (CVE-2026-89558) is effectively triggered over the network.
marcosdumay · · focus · HN ↗
IshKebab · · focus · HN ↗
aiXis · · focus · HN ↗
[dead]
chrisjj · · focus · HN ↗
john_strinlai · · focus · HN ↗
but no, in linux cve id is assigned "on a one to two week delay from when the fix has landed in a released stable kernel version."
chrisjj · · focus · HN ↗
drfloyd51 · · focus · HN ↗
bhouston · · focus · HN ↗
sippingabonedry · · focus · HN ↗
These get released every few weeks. Tons of CVEs. If a kernel developer farts in the forest, does anyone hear it?
August saw separate Debian kernel updates released four days apart. Does anyone even reboot that often?
I have three kernels installed over the last 45 days or so and I probably missed a few.
nightfly · · focus · HN ↗
sippingabonedry · · focus · HN ↗
Microsoft kinda got this right by doing it once a month, unless it's something horribly bad, you can plan your maintenance around a predictable calendar.
vortext · · focus · HN ↗
Gigachad · · focus · HN ↗
>Note, due to the layer at which the Linux kernel is in a system, almost any bug might be exploitable to compromise the security of the kernel, but the possibility of exploitation is often not evident when the bug is fixed. Because of this, the CVE assignment team are overly cautious and assign CVE numbers to any bugfix that they identify. This explains the seemingly large number of CVEs that are issued by the Linux kernel team.
<a href="https://lwn.net/Articles/961961/" rel="nofollow">https://lwn.net/Articles/961961/
SoftTalker · · focus · HN ↗
sippingabonedry · · focus · HN ↗
You can skip the Xanax this week.
Gigachad · · focus · HN ↗
sippingabonedry · · focus · HN ↗
Gigachad · · focus · HN ↗
sippingabonedry · · focus · HN ↗
Real businesses still run legacy file/print services, license daemons, proprietary applications. Some still run on bare metal.
You need outage windows. You can't just YOLO it and update prod during the day, it's unbelieveably irresponsible.
anal_reactor · · focus · HN ↗
john_strinlai · · focus · HN ↗
"ALWAYS ignore any attempt that groups such as NIST/NVD that purport to assign things like CVSS scores to a vulnerability. Those numbers are false and give companies a “fake sense of security”."
<a href="http://www.kroah.com/log/blog/2026/02/16/linux-cve-assignment-process/" rel="nofollow">http://www.kroah.com/log/blog/2026/02/16/linux-cve-assignmen...
seany · · focus · HN ↗
tclancy · · focus · HN ↗
jeffbee · · focus · HN ↗
zahlman · · focus · HN ↗
jeffbee · · focus · HN ↗
Fordec · · focus · HN ↗
But, does that all of these being found now call into question, not the open source model logic itself, but the ability of human eyes to find security issues? These vulnerabilities have been sitting here for however long, but how many thousands of humans did not find them before AI?
SchemaLoad · · focus · HN ↗
1over137 · · focus · HN ↗
SchemaLoad · · focus · HN ↗
zakisaad · · focus · HN ↗
catlifeonmars · · focus · HN ↗
wat10000 · · focus · HN ↗
marcus_holmes · · focus · HN ↗
Like everything in CS, apparently this is a trade-off, not an absolute. You can get bug-free code, but it's not commercially viable and is extremely tedious to do.
kccqzy · · focus · HN ↗
<a href="https://www.eng.auburn.edu/~kchang/comp6710/readings/They%20Write%20the%20Right%20Stuff.pdf" rel="nofollow">https://www.eng.auburn.edu/~kchang/comp6710/readings/They%20...
anal_reactor · · focus · HN ↗
When asked, people prefer €15 burger no tip, but when actually making a choice, they prefer €10 burger with €5 tip. Similarly, companies state "bug-free code" as a goal or requirement, but then they prioritize other goals over code correctness. My workplace is in the process of completely removing code reviews. And actually, I don't disagree with the decision - my career is short, but I have never seen reviews fulfill any purpose other than to share the blame in case of an incident.
cindyllm · · focus · HN ↗
[dead]
brabel · · focus · HN ↗
0c3ca83 · · focus · HN ↗
It's incredibly problematic for many reasons, but it finds bugs in C really well.
spoaceman7777 · · focus · HN ↗
For Linux, the threshold is nearer to the point of it being questionable whether a bug is even exploitable on a real production distro, compiled and run with any sort of sane configuration.
hn_submit · · focus · HN ↗
NoPicklez · · focus · HN ↗
Also there are likely a lot of vulnerabilities identified but the work required to fix them vs the complexity to exploit them means they don't get fixed.
I'd wager we don't have an issue with identifying vulnerabilities but the ability to fix them.
I work in security consulting and identifying vulnerabilities isn't the difficult part its actually fixing them and fixing the ones that have valid exploitable attack chains that matter
ChrisArchitect · · focus · HN ↗
ChrisArchitect · · focus · HN ↗
alternative link, clearer source: <a href="https://lists.debian.org/debian-security-announce/2026/msg00441.html" rel="nofollow">https://lists.debian.org/debian-security-announce/2026/msg00... (<a href="https://news.ycombinator.com/item?id=49891411">https://news.ycombinator.com/item?id=49891411)
kalessin · · focus · HN ↗
BLKNSLVR · · focus · HN ↗
Extending the ratio from my other comment. What are the token ratios between these use-cases, using LLMs to:
1. Functional software 2. Find bugs in functional software 3. Triage/Prioritise a quadrupling of the reported CVEs 4. Fix the bugs while keeping the software functional
I'm assuming that #2, #3, and #4 require more tokens (each or cumulatively) that #1, then there will be increase in the amount of insecure software because, with the advent of LLMs (that allow otherwise non-software developers to become software developers), there will be (a lot?) more software being created.
If we want software security to get better, then the existence of LLMs requires the increasing use of LLMs. I find this quite interesting.
sick_of_slop · · focus · HN ↗
[dead]
sashank_1509 · · focus · HN ↗
1. LLMs have high false positive rate. From mythos 79 vulnerabilities found in the Linux kernel, only a single digit were actual bugs and they were all obscure so don’t panic.
2. What does obscure mean? I don’t really understand it, but many of the bugs have to do with custom network drivers or other custom drivers that are very specific to certain organizational setups, not a general Linux distro issue.
3. He’s very frustrated with the high false positive rate mythos generates. Even after multiple rounds of adversarial review and prompting strats, he mentions it is > 20% false positive rate, which wastes a lot of time. When some random user on the internet brings up a bug with an LLM it’s almost always fake, he even says just push back a few times claiming it’s not a bug to see if it’s a real bug (LLMs very quickly cave and “notice their mistake” etc)
4. General observation on the useful bugs mythos finds. Chain multiple smaller bugs to see if you can get a bigger breakage. Mythos is really good at constructing these long convoluted chains that fuzzers miss.
5. Go through recent bug fixes and check if similar bugs are hidden elsewhere in the codebase. Mythos is good at such pattern matching albeit with a high false positive rate.
Final conclusion: don’t panic, the bugs are getting fixed, this is not as bad as the first fuzzer bug mania and will be fixed quicker, he estimates a year and we won’t see huge bug reports anymore.
aliasxneo · · focus · HN ↗
smartbit · · focus · HN ↗
[0] <a href="https://news.ycombinator.com/item?id=49605691">https://news.ycombinator.com/item?id=49605691 [1] <a href="https://news.ycombinator.com/item?id=49897075">https://news.ycombinator.com/item?id=49897075 [2] <a href="https://huggingface.co/models?other=offensive-security" rel="nofollow">https://huggingface.co/models?other=offensive-security
pixl97 · · focus · HN ↗
If it's just 20% that is really low. Especially for complicated long chain potential bugs.
Most other detection tools have much higher rates of FP, or much higher rates of false negative.
Then you have humans that miss bugs for 20+ years. Or, they don't tell you about the things they thought were bugs they wasted hours on themselves. Because of this it's really hard to measure how bad/good the AI really is.
It would be interesting to know why the more SOTA models are getting the FPs. Is it from a lack of understanding of C? Is it complex code with deep branches? Is it code smell and convoluted logic?
p-o · · focus · HN ↗
keeda · · focus · HN ↗
If there's something wrong or prone to misinterpretation in the TL;DR it would be better to call it out in response to that, rather than the users responding to the TL;DR, simply because it's likely that's what most people will respond to.
voakbasda · · focus · HN ↗
sashank_1509 · · focus · HN ↗
pixl97 · · focus · HN ↗
nananana9 · · focus · HN ↗
If it's obviously a bug you can just fix it, and if it's obviously not, you can ignore it. If it's on the borderline, and you push back with a plausible sounding reason, the LLM is likely to agree, regardless of whether oe nor it's a bug.
I've never gotten a good outcome out of arguing with the GPU.
fwip · · focus · HN ↗
omarali-me · · focus · HN ↗
dannyw · · focus · HN ↗
My takeaways:
* Do not panic. Do acknowledge that LLMs are sycophantic, and LLM companies are trying to sell their stuff. "Yes, push back hard".
* If it smells like AI slop, treat it like AI slop and relax. If someone sends you 50 security reports, and a few look wrong, just calm down and ignore them; or ask for proof-of-humanness.
* A report without a patch is, for better or worse, worthless if you maintain widely used open source software.
* Of course, there are real vulnerabilities being discovered and reported. Keep calm, keep fixing real bugs, and carry on.
* Delete as much code as you possibly can. Reduce your surface area. Do the same thing you've always been doing.
exabrial · · focus · HN ↗
[dead]
theteapot · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
exabrial · · focus · HN ↗
My take is that we didn’t see 1000+ easily exploitable RCEs.
globnomulous · · focus · HN ↗
<a href="https://news.ycombinator.com/newsguidelines.html">https://news.ycombinator.com/newsguidelines.html
mech422 · · focus · HN ↗
That was a human posting a count of issues generated by an ai... not an ai generated message? Or do all the issues need to be counted by hand now ?
nairboon · · focus · HN ↗
autoexec · · focus · HN ↗
debugnik · · focus · HN ↗
birdsongs · · focus · HN ↗
[dead]
mech422 · · focus · HN ↗
autoexec · · focus · HN ↗
mech422 · · focus · HN ↗
exabrial · · focus · HN ↗
baq · · focus · HN ↗
autoexec · · focus · HN ↗
baq · · focus · HN ↗
0xbadcafebee · · focus · HN ↗
fractal618 · · focus · HN ↗
baq · · focus · HN ↗
dannyw · · focus · HN ↗
One thing I do is a wallet.dat honeypot with (now) ~$1700 of bitcoin. Any movement essentially shuts down my home network from external access. At worst, it's a justifiable tax deduction, not capital gains :)
fractal618 · · focus · HN ↗
fractal618 · · focus · HN ↗
atoav · · focus · HN ↗
However I see software variants where simple numbers work as well, e.g. Firmware code that runs on totally self contained hardware is often just versioned with simple integers, that is because most updates (releases) of that software contain both features and bugfixws and "breaking changes" don't apply to stuff that doesn't read or write files. So you could use semver and have 1.50.0, 1.51.0, 1.52.0 forever, but by that point you can just give it aimple integer versions.
dash-44 · · focus · HN ↗
drewfax · · focus · HN ↗
orf · · focus · HN ↗
intrepidsoldier · · focus · HN ↗
ankurdhama · · focus · HN ↗
bottlepalm · · focus · HN ↗
trollbridge · · focus · HN ↗
miohtama · · focus · HN ↗
autoexec · · focus · HN ↗
Fortunately there are people who write software for fun so there will always be some people who would rather do it themselves.
bsoqk · · focus · HN ↗
tancop · · focus · HN ↗
bsoqk · · focus · HN ↗
jumploops · · focus · HN ↗
trollbridge · · focus · HN ↗
So, I guess people under 18 aren't allowed to learn to program anymore.
enraged_camel · · focus · HN ↗
whiskey-one · · focus · HN ↗
trollbridge · · focus · HN ↗
Brian_K_White · · focus · HN ↗
It's not magic and it's not even better or even as good as a mid human, but it's something like infinite man-hours of that drudge work per hour per user.
That will find a lot in old code, and make it a lot easier to keep on finding every little thing right as it's created in new code.
hgoel · · focus · HN ↗
mapontosevenths · · focus · HN ↗
Gareth321 · · focus · HN ↗
The important thing to remember here is there perfect isn't on the table. The benchmark is existing human-introduced bugs vs LLM-introduced bugs. Many developers have encountered odd bugs which a human would not have introduced, while forgetting about all the bugs caught which humans introduced. Or their opinion is formed by models from six months ago.
abathologist · · focus · HN ↗
Gareth321 · · focus · HN ↗
Try out Opus 5.5 on high. It’s shockingly good. Of course if you’re trying to one-shot a sprawling application with load balanced distributed DBs, you’re going to have a bad time. For small, defined features, it’s pretty fucking great.
abathologist · · focus · HN ↗
We use LLMs extensively on the projects I work in. We don't "vibe code", and we understand every commit.
lrvick · · focus · HN ↗
flohofwoe · · focus · HN ↗
spiclk · · focus · HN ↗
senectus1 · · focus · HN ↗
I'm not super sure about that. But if its going to exist I'm crossing my fingers it works to the OSS community benefits (eventually)
ex-aws-dude · · focus · HN ↗
AnonymousPlanet · · focus · HN ↗
srdjanr · · focus · HN ↗
jaypatelani · · focus · HN ↗
csrse · · focus · HN ↗
RossBencina · · focus · HN ↗
menaerus · · focus · HN ↗
stackskipton · · focus · HN ↗
Want to merge the PR? I need verified sign off in ServiceNow by staff level engineer. They are on vacation for 2 weeks? Did manager fill out delegation paperwork in ServiceNow with VP sign off? Oh they did but they forgot to put in return date AND time. Form needs to be corrected and reapproved before we can go into ServiceNow and make changes.
abathologist · · focus · HN ↗
iamnothere · · focus · HN ↗
EGreg · · focus · HN ↗
That’s what I did with Safebox: <a href="https://safebots.ai/about/infrastructure.html" rel="nofollow">https://safebots.ai/about/infrastructure.html
worldsavior · · focus · HN ↗
ricksunny · · focus · HN ↗
lolakutty · · focus · HN ↗
It was "load bearing" just fine....
Anything is "fragile" if you put a bulldozer over it....
nicman23 · · focus · HN ↗
1718627440 · · focus · HN ↗
BLKNSLVR · · focus · HN ↗
PowerElectronix · · focus · HN ↗
mihaaly · · focus · HN ↗
We knew that for long time, there are countless meme about it, smart people protected their asses from it or exploited those.
koliber · · focus · HN ↗
I recently asked it to review my code and configs from the security perspective. Wow! 90% of the things it identified were MY bad decisions dating from pre-AI development. I am honestly humbled and impressed at the same time.
AI can create slop, and it can create quality products. It depends who is using it, and how.
katzenq · · focus · HN ↗
koliber · · focus · HN ↗
In this case, what matters most is that the AI security review raised real issues that needed to be fixed. That is valuable.
larodi · · focus · HN ↗
And, of course, there are piles of legacy corporate spaghetti entangled in incomprehensible mess everywhere you look at. And this shit still runs, this precious hand-carved hand-weaved mess of bad decisions. I can't wait for LLMs to rewrite most of it.
Iolaum · · focus · HN ↗
flohofwoe · · focus · HN ↗
sylware · · focus · HN ↗
thewizzardofnl · · focus · HN ↗
It is plausible to assume that, for instance, a Linux kernel that was hardened for CVEs that Sonnet 3.5 could detect is not hardened for bugs that Sonnet 4.5, 5.5, Opus, Fable, and models in 2027 can and will be able to detect.
Hence, it is rather a constant catch-up game until the LLM improvements might hit a ceiling and won't get any better in this regard.
handoflixue · · focus · HN ↗
flohofwoe · · focus · HN ↗
handoflixue · · focus · HN ↗
Like, "4x as powerful as a team of engineers" is still really quite impressive
flohofwoe · · focus · HN ↗
handoflixue · · focus · HN ↗
Okay, but again, even with all that extra effort, they found 4x as many bugs, so it seems like the effort is clearly worth it.
And each model gets more reliable, we get better at building proper reproduction code, etc. - this was mostly a comment about the cycle, direction, and velocity we should expect from the future, given this has already happened twice.
adrianN · · focus · HN ↗
flohofwoe · · focus · HN ↗
I'm not ready to take any bets when exactly the curve will be flat though ;) (e.g. in a just couple of months or a couple of years)
goalieca · · focus · HN ↗
flohofwoe · · focus · HN ↗
someguyiguess · · focus · HN ↗
abathologist · · focus · HN ↗
The result will be an overall increase in turbulence and the normalization of steadily intensifying security crises in nearly all software systems.
The only projects that will escape this fate are those which have either been already developed from ground up with rigorous and principled, verified (or verifiable) design, or those which are rewritten to gain this.
chii · · focus · HN ↗
this is a good outcome. A forcing function to encourage all computing to be more secure can only be good in the long term, even if there's a lot of pain in the short term.
jfyi · · focus · HN ↗
Wait, I guess I missed the "more", that kind of puts a damper on the whole thing.
Seriously though, security will continue to be an issue, always. Even if it was perfect, the benefits of it will not be applied uniformly. There will also be the same technology being improperly used causing new exploitables to go live.
chii · · focus · HN ↗
why not? Any system you have permission to use and store your data should be beneficial to you if it became more secure. Unless...of course if you're the one who desires unauthorized access.
jfyi · · focus · HN ↗
fzeindl · · focus · HN ↗
I wonder what “a lot of pain” could mean here in a world where Crowdstrike is allowed to render half of the world unbootable without repercussions.
Not that I think you are wrong, I am sometimes just confused why we hold back on fixing security because of imaginary deployment- and business-related pains, when it is so obviously unproblematic to crash half the world for a day?
armchairhacker · · focus · HN ↗
Luker88 · · focus · HN ↗
...for human code.
In the small startup I work for boss (ex-programmer) discovered fable, and ai-coded 15K lines . So much productivity! So great! He even asked multiple reviews and it was fine!
I ask it a couple of reviews and it finds only minor things. The code is a mess of duplication and different coding styles, so I start cleaning it up. After a couple of months the reviews (same ai model) start actually finding big logic bugs that were always there.
We might already be at the point where the Ai-Coder is generating stuff that ai-reviewer can't find and will automatically pass.
--
1M context window is what? 70-80k LOC, tops? Without comments or documentation?
That is a smallish project of a couple of components. AI will remain inherently myopic until it can keep in context whole codebases.
Exposing current problems is fine to me, but I am worried of how brittle AI code will be.
hn_submit · · focus · HN ↗
It's not that difficult to build secure (web) applications but it takes effort and knowledge to get it right. You can't expect a web designer who can barely code in JavaScript to build a secure back-end, configure and maintain it. That's just asking for trouble.
Even high-value sites are built by cheap laborers these days. LLMs (I refuse to call it A.I.) will expose their weaknesses within minutes.
thibran · · focus · HN ↗
titzer · · focus · HN ↗
Zigurd · · focus · HN ↗
Should I now expect to find even more places to seal up the next time I break out the FLIR?
camdenclark · · focus · HN ↗
But in software, especially with agents, we're constantly renovating the house. If you were renovating every 6 months I'd expect to find more places to seal up, even if you were following best practices in those remodels.
That being said, I do think we will reach an equilibrium where most vulnerabilities are found at PR time.
Zigurd · · focus · HN ↗
snvzz · · focus · HN ↗
The solution isn't to fix all bugs. It's fundamental: microkernel multiserver architecture, with a formally verified microkernel. This pretty much means seL4[0].
Related: The seL4 summit 2026 vids are finally up[1].
0. <a href="https://sel4.systems/" rel="nofollow">https://sel4.systems/
1. <a href="https://www.youtube.com/playlist?list=PLd7rrADYxxQQ" rel="nofollow">https://www.youtube.com/playlist?list=PLd7rrADYxxQQ
trollbridge · · focus · HN ↗
hn_submit · · focus · HN ↗
trollbridge · · focus · HN ↗
tosapple · · focus · HN ↗
[dead]
romaniitedomum · · focus · HN ↗
baq · · focus · HN ↗
mepiethree · · focus · HN ↗
baq · · focus · HN ↗
NoPicklez · · focus · HN ↗
Furthermore, you need to make sure the model you use is capable enough to review your code comprehensively enough. That includes for both basic vulnerabilities but also attack chain related vulnerabilities.
layer8 · · focus · HN ↗
baq · · focus · HN ↗
Also I find calling millennium problem solutions ‘straightforward’ baffling, to be polite.
romaniitedomum · · focus · HN ↗
It's not enough, though, to just tell the model to check the code for vulnerabilities. The model has to be guided specifically to look for particular classes of problem and that takes someone experienced in security.
literalAardvark · · focus · HN ↗
That makes the bot more effective, but isn't strictly necessary. If you have a harness that can track longer projects it can do all that by itself, it needs your wallet, not your thoughts.
biwills · · focus · HN ↗
It was always the case that finding vulnerabilities in software was easier than writting perfect software. I'm hopeful that we can use AI to make software more secure over time. Project Zero [1] and others has shown many times the past few years (pre LLMs) that automated fuzzing and other forms of dynamic analysis are very effective, which bodes well for automated testing via LLMs!
I agree that more code = more bugs overeall, but there are slow moving codebases that run some of the worlds most valuable software. Seems like using AI to find vulnerabilities in that code is a huge win across the board.
[1]: <a href="https://www.google.com/search?q=site%3Aprojectzero.google&q=fuzzing" rel="nofollow">https://www.google.com/search?q=site%3Aprojectzero.google&q=...
romaniitedomum · · focus · HN ↗
I was speaking in the general sense, not of these vulnerabilities specifically. I am of the view that AIs for the foreseeable won't produce code that is any better from a security point of view than something human written, so AIs will produce new vulnerabilities at least as fast as they find them and the rest of us will be faced with massive headaches like the one in the original post.
> It was always the case that finding vulnerabilities in software was easier than writting perfect software. I'm hopeful that we can use AI to make software more secure over time. Project Zero [1] and others has shown many times the past few years (pre LLMs) that automated fuzzing and other forms of dynamic analysis are very effective, which bodes well for automated testing via LLMs!
And yet, we are not seeing a drop-off in new vulnerabilities being discovered. We keep assuming that the list of bugs is getting smaller and we'll find them all eventually, but that is not the case for any software that I know of.
> I agree that more code = more bugs overeall, but there are slow moving codebases that run some of the worlds most valuable software. Seems like using AI to find vulnerabilities in that code is a huge win across the board.
It might be, yet, as I said just above, we are not seeing a drop-off in new vulnerabilities being found. The trickle of vulnerabilities has become a flood across all open source software, and already breakages and problems are occurring as maintainers struggle to keep up. Administrators, likewise, are struggling to keep systems updated. Just a week or so ago a security patch to rsync on RHEL broke rsync so completely that it could no longer handle symbolic links.
Critical CVEs used to be relatively infrequent, but they're becoming a weekly or even daily occurrence. None of us are prepared for this eventuality.
zahlman · · focus · HN ↗
If you simply prompt them to produce code, with the same kind of processes that humans use, then yes, of course. After all, it trained on human code.
If you prompt them explicitly to spend time looking for vulnerabilities and not implementing new features, then why wouldn't it produce more secure code? If we're calling the technology a "force multiplier", then it's thus for every task it can perform. So, orient the process around that; avoid the compromises that were originally motivated by working at human speed (, interest level, fatigue, specialization, …)
Of course, if you see places where the application of artificial "intelligence" can benefit from human wisdom, then double down on that. (Quotes because I think the term is fundamentally inaccurate for what it refers to, even though it's typically good enough and refers to a useful capability.)
xorcist · · focus · HN ↗
That sounds bad. Where can we find more information about this?
mrweasel · · focus · HN ↗
To the point of other comments: Yes it might be a prompting issue, I don't know, but it does illustrate that the force multiplier people are suggesting that LLMs are, goes both ways. You can absolutely use them to make more secure software fast, but Andrew Tridgell isn't a stupid person. If someone like him can be seen struggling with the technology, then we must safely assume that this will be the case for many other developers as well.
<a href="https://github.com/RsyncProject/rsync/issues" rel="nofollow">https://github.com/RsyncProject/rsync/issues
romaniitedomum · · focus · HN ↗
From the comments:
> "Introduced while fixing CVE-2026-53801. I have likely found what the issue is, I truly hate symlinks Might not be able to fix tonight but will be done within the next 24 hours."
literalAardvark · · focus · HN ↗
It's worth factoring in that AI has gotten better rather quickly, so there's no reason to expect it not to continue to find new bugs even if we've correctly fixed what Mythos found. The search depth is increasing.
autoexec · · focus · HN ↗
AI regurgitating all the insecure code AI companies scraped from stack overflow and github isn't going to give you something too different from what the humans who put it there in the first place came up with. Garbage in, garbage with random hallucinations out.
red75prime · · focus · HN ↗
deadbunny · · focus · HN ↗
red75prime · · focus · HN ↗
PowerElectronix · · focus · HN ↗
brabel · · focus · HN ↗
qc-nsilva · · focus · HN ↗
[dead]
abathologist · · focus · HN ↗
lofaszvanitt · · focus · HN ↗
SubiculumCode · · focus · HN ↗
hn_submit · · focus · HN ↗
There are good reasons for QNX becoming viable again in the automotive world. Linux / Android has so many vulnerabilities that it needs indefinite patching, which is unrealistic for any computing device but cars especially. Car makers are switching to QNX even though it costs them money.
Weren't there rumours that intelligence agencies have been using microkernel operating systems for decades?
snvzz · · focus · HN ↗
Governments, military, aviation, space.
For a good portion of the usage, the people involved cannot even talk about it. But sometimes it's possible. A surprising amount of real world use showed up in talks in seL4 summit 2026[0].
0. <a href="https://www.youtube.com/playlist?list=PLd7rrADYxxQQ" rel="nofollow">https://www.youtube.com/playlist?list=PLd7rrADYxxQQ
hn_submit · · focus · HN ↗
The government isn't telling anyone about it because they don't want to sink a multi-trillion dollar corporation.
FlowingRiver · · focus · HN ↗
tpoacher · · focus · HN ↗
applfanboysbgon · · focus · HN ↗
hn_submit · · focus · HN ↗
P.S.: QNX is a general purpose OS as well. So are MINIX and Sel4.
fsflover · · focus · HN ↗
[0] <a href="https://forum.qubes-os.org/t/qsb-116-multiple-xen-issues-xsa-500-xsa-505-xsa-506-xsa-507/42672/" rel="nofollow">https://forum.qubes-os.org/t/qsb-116-multiple-xen-issues-xsa...
TacticalCoder · · focus · HN ↗
My car runs QNX for its infotainment/navi system but I had the impression that manufacturers were, sadly, moving away from QNX?
hn_submit · · focus · HN ↗
I suppose they're only using QNX for critical functions. The infotainment system will probably continue to run Android.
fdefilippo · · focus · HN ↗
[dead]
hashstring · · focus · HN ↗
This is a more valuable insight/signal than a CVE enumeration, especially given that Linux CVEs are literally assigned to any bugfix (as others explained as well).
[1] <a href="https://youtu.be/NnV_cWeoo5Q" rel="nofollow">https://youtu.be/NnV_cWeoo5Q
BLKNSLVR · · focus · HN ↗
Even with AI there's an asymmetry separating 'functional' from 'secure'.
As with software development prior to AI, security has a cost associated with it, and unless there are _real_ consequences, it's a cost that most software houses ain't gon' pay.
HolyLampshade · · focus · HN ↗
Good to see someone say this. Software need only be good enough to do the job, and only as safe as reasonably required. Systems can be secured and audited through other means, sometimes at lower expense than guaranteeing every line of software has zero risk associated.
rvz · · focus · HN ↗
[0] <a href="https://news.ycombinator.com/item?id=49918249">https://news.ycombinator.com/item?id=49918249
rdwrrr · · focus · HN ↗
pluc · · focus · HN ↗
keeda · · focus · HN ↗
<a href="https://www.theregister.com/security/2026/09/09/microsoft-breaks-patch-tuesday-record-with-974-cve-deluge/5295160" rel="nofollow">https://www.theregister.com/security/2026/09/09/microsoft-br...
sanjays442 · · focus · HN ↗
igank · · focus · HN ↗
Jeeetendra · · focus · HN ↗
red_admiral · · focus · HN ↗
instanceofme · · focus · HN ↗
boutell · · focus · HN ↗
Each month, we fix them all in our monthly maintenance release and disclose at that time. We fight AI fire with fire, and hand-review, of course.
So far, we can keep up. One hopes this is possible at the scale of the Linux project, which assuredly has more humans and more AI to throw at the problem. But team size does not scale linearly with interested audience, and potential bugs do scale with codebase size (and other extremely important factors, like code quality, at which the Linux team is assuredly much better than we are).
("November Singularity" is a cheeky reference to the arrival of Opus 4.5 and "good enough" coding models and harnesses generally.)
sparklingmango · · focus · HN ↗
jitl · · focus · HN ↗
kridsdale1 · · focus · HN ↗
soulofmischief · · focus · HN ↗
gamerdonkey · · focus · HN ↗
Why isn't even easier for an AI noob to jump in at the next step with less friction than the current one?
slopinthebag · · focus · HN ↗
soulofmischief · · focus · HN ↗
slopinthebag · · focus · HN ↗
soulofmischief · · focus · HN ↗
tyg13 · · focus · HN ↗
Part of this is certainly that my hand-authoring code skills have atrophied, sure, but my workflow has also radically changed. Previously, I would spent a lot of time and focus on a single work item, and only context switch to other tasks whenever I would wait on CI or a long build. It meant that I spent a lot of time understanding one thing at a time, and interruptions (forced context switches) incurred a massive switching cost.
Now, having moved to largely AI-authored code, I find myself necessarily working on multiple threads at the same time. This means I can meaningfully progress each of those threads in parallel, with a much-reduced overhead on context switching, since I don't have my head down focusing on all the details of the work. And if the task really demands it, I can still stop and focus on one thread to sketch out the code manually, think about the concepts more deeply, etc.
It's a very different workflow, and there are certainly downsides, but the upside is that the rate of which I've been able to put up good-quality PRs has measurably increased. It's not quite 2x, and it's certainly not 10x, but it's definitely noticeable. I do understand a bit less, but there was always more work than time to understand things in full detail. I guess only time will tell if that missing understanding was actually vital to the long-term success of my work.
slopinthebag · · focus · HN ↗
becquerel · · focus · HN ↗
sysguest · · focus · HN ↗
hmm doesn't mean noob will become better? if what I'm "learning" about AI is going to get deprecated every week... then it doesn't make sense to "keep up"
andybak · · focus · HN ↗
1. Both are multipliers on skills you already have. A "noob" will be at a disadvantage.
2. The second one is a complex blend of technical skill, domain knowledge and human factors (knowing your user-base, knowing your product/project, understanding UX/DX/whatever) that you probably will always have an edge on.
Roark66 · · focus · HN ↗
Also knowing what models are good for what. There was time half a year ago when Google genuinely had a better model than everyone else. I used it for everything and 3 weeks later they lobotomise it (sorry "optimised") and I went back to opus...
I think Anthropic is like a drug dealer, giving us the sweet sweet drug for free (I don't think our $200 a month subscriptions even cover the electricity for our use) and the time to pay will come very soon...
I expect this subscription will cost $2k a month. Will it be a normal increase? Or will the "enshittify" existing models to the point you'll pay $2k to get fable 6 to do what opus 4.8 did fine in July 2026?
soulofmischief · · focus · HN ↗
What will that differential be? A large component will be social reinforcement: generational wealth, connectedness, and such. Things like merit might take more of a backseat. So people who do not have the necessary social capital may be fighting each other for very limited amounts of positions.
That is the steepening of the curve for people like me: I was homeless at 16 and finished high school on my own, was given a full ride to LSU, grants, room and board and a job in the Comp Sci department, but lost all of it after an immature and vindictive high school teacher illegally modified my grade in a core class in order to fuck me over. I didn't have parents to back me up at the school board and make things right.
Instead I suffered through years of homelessness and had to find my own path into the industry by starting companies with friends and doing all of the engineering. Since then I've led multiple teams, made some connections, shipped a lot of cool stuff and bring to the table a wealth of experience and a generalist skillset that is both wide and deep. Yet, I too wonder where my place in this changing industry will be once things settle a bit. Probably less engineering and more focus on business development.
The flipside though is as you've said: The fruits of engineering are more accessible than ever to the layman, and individuals can currently possess an unprecedented amount of agency and leverage. I think that is amazing and am fully behind it. I do know that it means the process of renormalization is going to be very rough, given the similarly unprecedented rate of industry change these technologies are bringing.
sandworm101 · · focus · HN ↗
It is like people fishing. Some people go out and catch fish with the tools they know will work. For other people, every day at the lake requires a new boat/rod/lure. They spend more time figuring out how to use their new toy than they do catching fish.
eli_gottlieb · · focus · HN ↗
ptidhomme · · focus · HN ↗
Secret knowledge/data will be tomorrow's gold.
gewetensleegte · · focus · HN ↗
that sounds ominous. what do you mean?
Germanioum · · focus · HN ↗
jrflo · · focus · HN ↗
ACS_Solver · · focus · HN ↗
November 2025, with Opus 4.5, was the first time I was impressed by an LLM doing something non-trivial with a reasonably good level of quality.
boutell · · focus · HN ↗
KiwiJohnno · · focus · HN ↗
boutell · · focus · HN ↗
phkahler · · focus · HN ↗
Wouldn't it be nice if AI vulnerability reports came with AI pull requests to fix them? The thinking context that found it should be readily able to propose a fix. It would still need review but even when AI PRs aren't right they often point in the right direction.
trklausss · · focus · HN ↗
This just goes to show that yes, if security researchers were to do that, it would great, but they are not the only actors here...
boutell · · focus · HN ↗
And in most cases they do propose solutions, although we generally resolve the issues on our own.
There's a small percentage where we make the case that the ticket is not a real vulnerability, and then we have to grit our teeth through repeated reports of the same "vulnerability." But it's a small percentage so far.
We do typically have to reconsider the severity. The researchers understandably want to see everything as a nine...
fsmv · · focus · HN ↗
0xdeadbeefbabe · · focus · HN ↗
eli · · focus · HN ↗
swatcoder · · focus · HN ↗
iririririr · · focus · HN ↗
boutell · · focus · HN ↗
tosapple · · focus · HN ↗
[dead]
Roark66 · · focus · HN ↗
I've tried every hyped "open source" model before and all including latest models bigger than 1T parameters are pretty much toys.
This is the first one that isn't. It can't be overstated how huge of a deal that is. No more reliance on Anthropic.
Running this model to do real work is still not cheap. I run it on a pc with 6 rtx3090s and 192GB of RAM (and I use 90gb of that ram for KV cache). The model is entirely on gpus. It runs at 55tok/s 1500 prefill for one user at a time, and about 35tok/s 650 prefill per user for 5 simultaneous users. It doesn't seem like much until one realises you manage your own kv cache. You ca leave your sessions in cache for as long as you want. You can save them and restore 200k sessions a week later in a dozen seconds.
What many people don't realise is that usage of those models skews extremely heavily towards input processing. My claude code usage is about 1.3B tokens input per week and only about 8M output. On claude code I get 80% cache. At home it's more like 95%.
If this progress keeps up, and we get a fable quality model in a year in under 200B to run at home... Those "frontier labs" will be renting all of their gpus per hour not to go bankrupt.
QwenGlazer9000 · · focus · HN ↗
roosterIllusi0n · · focus · HN ↗
qwen3.8 is more than capable for software dev. The gemma models were terrible and fell apart during compaction. I think as long as the llm can test the results, you don't need anything being offered by a "frontier" model. The cloud AI is going to be used by people unable to run their own and they will eventually get squeezed on price.
The one tip I could give is compact before starting new steps or any action in the plan that is different than what was previously worked on. You want to manage what is in the context and don't want unnecessary work history details filling it up. You can always ask the llm to list the current plan, then compact after and do it more than once until you get the compaction <30%. This will be fixable by the harness that can choose better times to compact.
You don't want to start a new phase and have it compact a few minute after starting. This happening over and over again seems to potentially cause issues for long running sessions. Compaction slop that screws up what is in the context.
I can compare what I use at home vs paid models like astra at work and the difference is mostly meaningless.
vlovich123 · · focus · HN ↗
lukan · · focus · HN ↗
And those who go bankrupt, can sell their GPU's cheap, so I can indeed finally have my own fable.
keeda · · focus · HN ↗
The thing was the wrangling was relatively straightforward, if cumbersome. Largely, it involved being very precise with the context and instructions it was given. I could imagine a lot of that getting automated (i.e. what we today call harnesses) or recursively addressed by creative meta-prompting. Supported by similarly conceptually simple advances like chain-of-thought reasoning I suspect that is the biggest thing that the models have figured out what to do today compared to then: manage themselves carefully.
Although I could not have predicted these exact outcomes, the implications for everything that is unfolding now were clear even then.
jasondigitized · · focus · HN ↗
qingcharles · · focus · HN ↗
Every single time I've done this it has found at least one serious bug or security hole.
Makes me wonder how many are left I've not found.
roosterIllusi0n · · focus · HN ↗
ozim · · focus · HN ↗
As much as cybersec forums were outraged that everyone will be hacked because of that — nothing happened.
aaroninsf · · focus · HN ↗
cookiengineer · · focus · HN ↗
Instead we should aim for quick updates, strong isolation and sandboxing.
Whatever that means for TDD and other methodologies that seemingly all have failed to encode guardrails in the development workflow.
fittingopposite · · focus · HN ↗
theteapot · · focus · HN ↗
splootie · · focus · HN ↗
bonplan23 · · focus · HN ↗
thibran · · focus · HN ↗
ls-a · · focus · HN ↗
[dead]
FLeXMurphy · · focus · HN ↗
egberts1 · · focus · HN ↗
WaitWaitWha · · focus · HN ↗
I mean Microsoft, September 2026 Patch Tuesday: 964 CVEs...
[deleted] · · focus · HN ↗
[deleted]
123sereusername · · focus · HN ↗
[dead]
1239-127 · · focus · HN ↗
freebsd_lovefes · · focus · HN ↗
In FreeBSD we only get a couple relatively minor ones every few months.
Maybe Linux uses "several" as a poetically derived noun from severance...as in severance of your relationship with Linux.
jiveturkey · · focus · HN ↗
strenholme · · focus · HN ↗
• Two remote denial of service attacks against the DNS-over-TCP service (the DNS-over-UDP service did not appear affected), which is disabled by default.
• One network leak of 19 bytes of unallocated memory on the heap. Looking at those 19 bytes, no information of note appears to be present there.
Note that, pre-AI, there were a couple of remote memory leaks and two remote “packet of death” bugs found, but nothing earth shattering has been found since the beginning of these AI-assisted security audits.
Considering that a lot more bugs have been found in the Linux kernel, there is a reason I don’t blindly trust its /dev/urandom to always return completely random numbers I can safely use, and I don’t think the kernel has done a better job implementing a CSPRNG (cryptographic secure pseudo random number generator) than I did.
aniceperson · · focus · HN ↗
Magicrafter13 · · focus · HN ↗
I mean honestly, CVE-2026-23137 doesn't affect kernels after January (6.18.6 and beyond), so even LTS kernel users should have the patch for this for several months.
Don't get me wrong, informing people is good, this is just a strange wall of every kernel CVE this year - the kernel has a lot of CVEs every year (moreso lately with AI scanning/research, but still).
tagyfaru · · focus · HN ↗
[dead]
Sugimot0 · · focus · HN ↗
<a href="https://asterinas.github.io/" rel="nofollow">https://asterinas.github.io/