Several vulnerabilities have been discovered in the Linux kernel
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Several vulnerabilities have been discovered in the Linux kernel
Unofficial Hacker News client; not affiliated with Y Combinator.
intrepidsoldier · · focus · HN ↗
ankurdhama · · focus · HN ↗
bottlepalm · · focus · HN ↗
trollbridge · · focus · HN ↗
miohtama · · focus · HN ↗
autoexec · · focus · HN ↗
Fortunately there are people who write software for fun so there will always be some people who would rather do it themselves.
bsoqk · · focus · HN ↗
tancop · · focus · HN ↗
bsoqk · · focus · HN ↗
jumploops · · focus · HN ↗
trollbridge · · focus · HN ↗
So, I guess people under 18 aren't allowed to learn to program anymore.
enraged_camel · · focus · HN ↗
whiskey-one · · focus · HN ↗
trollbridge · · focus · HN ↗
Brian_K_White · · focus · HN ↗
It's not magic and it's not even better or even as good as a mid human, but it's something like infinite man-hours of that drudge work per hour per user.
That will find a lot in old code, and make it a lot easier to keep on finding every little thing right as it's created in new code.
hgoel · · focus · HN ↗
mapontosevenths · · focus · HN ↗
Gareth321 · · focus · HN ↗
The important thing to remember here is there perfect isn't on the table. The benchmark is existing human-introduced bugs vs LLM-introduced bugs. Many developers have encountered odd bugs which a human would not have introduced, while forgetting about all the bugs caught which humans introduced. Or their opinion is formed by models from six months ago.
abathologist · · focus · HN ↗
Gareth321 · · focus · HN ↗
Try out Opus 5.5 on high. It’s shockingly good. Of course if you’re trying to one-shot a sprawling application with load balanced distributed DBs, you’re going to have a bad time. For small, defined features, it’s pretty fucking great.
abathologist · · focus · HN ↗
We use LLMs extensively on the projects I work in. We don't "vibe code", and we understand every commit.
lrvick · · focus · HN ↗
flohofwoe · · focus · HN ↗
They're definitely a useful additional tool for finding more (and more obscure) bugs, but that takes a lot of both human and compute effort too (quite a few of the reported bugs are actually false positives on close inspection, and apparently even with the latest locked down "wonder weapon" models like Mythos), and after all the reports are clean and validated you still can't be 100% sure (but at least a bit more confident) that the code is now free of bugs.
spiclk · · focus · HN ↗
senectus1 · · focus · HN ↗
I'm not super sure about that. But if its going to exist I'm crossing my fingers it works to the OSS community benefits (eventually)
ex-aws-dude · · focus · HN ↗
AnonymousPlanet · · focus · HN ↗
srdjanr · · focus · HN ↗
jaypatelani · · focus · HN ↗
csrse · · focus · HN ↗
RossBencina · · focus · HN ↗
That may be true. Serious question though: even if most devs wanted to develop formally verified code, do you think that it is reasonable to suggest that the typical systems developer could do it with today's tools? I don't mean verified protocols (TLA+) or verified algorithms (SPIN) I mean end-to-end verified code, a-la seL4. I got the impression that this is still very specialised work. Perhaps things have advanced since I last checked.
menaerus · · focus · HN ↗
stackskipton · · focus · HN ↗
Want to merge the PR? I need verified sign off in ServiceNow by staff level engineer. They are on vacation for 2 weeks? Did manager fill out delegation paperwork in ServiceNow with VP sign off? Oh they did but they forgot to put in return date AND time. Form needs to be corrected and reapproved before we can go into ServiceNow and make changes.
abathologist · · focus · HN ↗
iamnothere · · focus · HN ↗
As it evolves I suspect there will be a push to verify more components of the stack. Once the capabilities layer can be verified, verification of most other components and drivers would become much less urgent.
EGreg · · focus · HN ↗
That’s what I did with Safebox: <a href="https://safebots.ai/about/infrastructure.html" rel="nofollow">https://safebots.ai/about/infrastructure.html
worldsavior · · focus · HN ↗
ricksunny · · focus · HN ↗
lolakutty · · focus · HN ↗
It was "load bearing" just fine....
Anything is "fragile" if you put a bulldozer over it....
nicman23 · · focus · HN ↗
1718627440 · · focus · HN ↗
BLKNSLVR · · focus · HN ↗
PowerElectronix · · focus · HN ↗
mihaaly · · focus · HN ↗
We knew that for long time, there are countless meme about it, smart people protected their asses from it or exploited those.
koliber · · focus · HN ↗
I recently asked it to review my code and configs from the security perspective. Wow! 90% of the things it identified were MY bad decisions dating from pre-AI development. I am honestly humbled and impressed at the same time.
AI can create slop, and it can create quality products. It depends who is using it, and how.
katzenq · · focus · HN ↗
koliber · · focus · HN ↗
In this case, what matters most is that the AI security review raised real issues that needed to be fixed. That is valuable.
larodi · · focus · HN ↗
And, of course, there are piles of legacy corporate spaghetti entangled in incomprehensible mess everywhere you look at. And this shit still runs, this precious hand-carved hand-weaved mess of bad decisions. I can't wait for LLMs to rewrite most of it.
Iolaum · · focus · HN ↗
flohofwoe · · focus · HN ↗
sylware · · focus · HN ↗
thewizzardofnl · · focus · HN ↗
It is plausible to assume that, for instance, a Linux kernel that was hardened for CVEs that Sonnet 3.5 could detect is not hardened for bugs that Sonnet 4.5, 5.5, Opus, Fable, and models in 2027 can and will be able to detect.
Hence, it is rather a constant catch-up game until the LLM improvements might hit a ceiling and won't get any better in this regard.
handoflixue · · focus · HN ↗
flohofwoe · · focus · HN ↗
handoflixue · · focus · HN ↗
Like, "4x as powerful as a team of engineers" is still really quite impressive
flohofwoe · · focus · HN ↗
In my hobby projects I use LLMs in my development workflow mainly for passive bug scanning, reviewing and helping to maintain tests, they are definitely useful for catching some bugs early and noticing unhandled edge cases, but they're also definitely no silver bullet (they sometimes ignore quite obvious bugs, and the fewer 'obvious' bugs remain the more one has to be careful about false positives). E.g. the funny thing is that now I'm actually slower than before due to the intense 'rubber ducking' with LLMs and cross-checking their results, but I still want to pretend that the resulting code is more robust out of the door.
Eg everything that Greg KH says in the video sounds very familiar, and it's very disappointing that Mythos still suffers from the same issues (or maybe even worse) as older models.
handoflixue · · focus · HN ↗
Okay, but again, even with all that extra effort, they found 4x as many bugs, so it seems like the effort is clearly worth it.
And each model gets more reliable, we get better at building proper reproduction code, etc. - this was mostly a comment about the cycle, direction, and velocity we should expect from the future, given this has already happened twice.
adrianN · · focus · HN ↗
flohofwoe · · focus · HN ↗
I'm not ready to take any bets when exactly the curve will be flat though ;) (e.g. in a just couple of months or a couple of years)
goalieca · · focus · HN ↗
flohofwoe · · focus · HN ↗
someguyiguess · · focus · HN ↗
abathologist · · focus · HN ↗
The result will be an overall increase in turbulence and the normalization of steadily intensifying security crises in nearly all software systems.
The only projects that will escape this fate are those which have either been already developed from ground up with rigorous and principled, verified (or verifiable) design, or those which are rewritten to gain this.
chii · · focus · HN ↗
this is a good outcome. A forcing function to encourage all computing to be more secure can only be good in the long term, even if there's a lot of pain in the short term.
jfyi · · focus · HN ↗
Wait, I guess I missed the "more", that kind of puts a damper on the whole thing.
Seriously though, security will continue to be an issue, always. Even if it was perfect, the benefits of it will not be applied uniformly. There will also be the same technology being improperly used causing new exploitables to go live.
chii · · focus · HN ↗
why not? Any system you have permission to use and store your data should be beneficial to you if it became more secure. Unless...of course if you're the one who desires unauthorized access.
jfyi · · focus · HN ↗
fzeindl · · focus · HN ↗
I wonder what “a lot of pain” could mean here in a world where Crowdstrike is allowed to render half of the world unbootable without repercussions.
Not that I think you are wrong, I am sometimes just confused why we hold back on fixing security because of imaginary deployment- and business-related pains, when it is so obviously unproblematic to crash half the world for a day?
armchairhacker · · focus · HN ↗
Luker88 · · focus · HN ↗
...for human code.
In the small startup I work for boss (ex-programmer) discovered fable, and ai-coded 15K lines . So much productivity! So great! He even asked multiple reviews and it was fine!
I ask it a couple of reviews and it finds only minor things. The code is a mess of duplication and different coding styles, so I start cleaning it up. After a couple of months the reviews (same ai model) start actually finding big logic bugs that were always there.
We might already be at the point where the Ai-Coder is generating stuff that ai-reviewer can't find and will automatically pass.
--
1M context window is what? 70-80k LOC, tops? Without comments or documentation?
That is a smallish project of a couple of components. AI will remain inherently myopic until it can keep in context whole codebases.
Exposing current problems is fine to me, but I am worried of how brittle AI code will be.
hn_submit · · focus · HN ↗
It's not that difficult to build secure (web) applications but it takes effort and knowledge to get it right. You can't expect a web designer who can barely code in JavaScript to build a secure back-end, configure and maintain it. That's just asking for trouble.
Even high-value sites are built by cheap laborers these days. LLMs (I refuse to call it A.I.) will expose their weaknesses within minutes.
thibran · · focus · HN ↗
titzer · · focus · HN ↗
Zigurd · · focus · HN ↗
Should I now expect to find even more places to seal up the next time I break out the FLIR?
Or you could categorize the current bug apocalypse under the heading "unsustainable trends will not be sustained."
camdenclark · · focus · HN ↗
But in software, especially with agents, we're constantly renovating the house. If you were renovating every 6 months I'd expect to find more places to seal up, even if you were following best practices in those remodels.
That being said, I do think we will reach an equilibrium where most vulnerabilities are found at PR time.
Zigurd · · focus · HN ↗