‹ BackHN Continuity

Thread

Several vulnerabilities have been discovered in the Linux kernel

576 points · 408 comments · luispa

  1. intrepidsoldier · · focus · HN ↗
    Just the beginning. AI is going to expose how fragile the entire computing infrastructure in our world is.
    1. ankurdhama · · focus · HN ↗
      Does this also mean the code generated and reviewed by LLMs will not have such issues going forward?
      1. bottlepalm · · focus · HN ↗
        It won't which means inevitably malicious AI will probably backdoor us. Damned if we do, damned if we don't.
      2. trollbridge · · focus · HN ↗
        Quite the opposite, particularly when the biggest vendors of coding agents insist on not allowing their models to be used to check the code they generate for security issues.
        1. miohtama · · focus · HN ↗
          Why stop there? We should regulate who is allowed to write code in the first place!
          1. autoexec · · focus · HN ↗
            Easily done if everyone can be convinced that learning to write code is pointless since you can just pay an AI company for access to a chatbot that will write it for you.

            Fortunately there are people who write software for fun so there will always be some people who would rather do it themselves.

          2. bsoqk · · focus · HN ↗
            So, like professional orders in Europe? Thankfully everybody agreed that writing code is not engineering so this isn't mandated by law, but we were this close.
            1. tancop · · focus · HN ↗
              You don't need a permit or degree to draw up plans, just to write your name on the official version and get it implemented in the physical world. It's closer to deploying code than writing.
              1. bsoqk · · focus · HN ↗
                We have a lot of that in Europe. Parasite professions. Someone who does nothing but charges a lot for his signature.
          3. jumploops · · focus · HN ↗
            “And that was how CS became a real engineering degree”
          4. trollbridge · · focus · HN ↗
            They sort of already do; the commercial American providers all require you to be 18 to sign up for a plan capable of agentic coding.

            So, I guess people under 18 aren't allowed to learn to program anymore.

        2. enraged_camel · · focus · HN ↗
          I've been able to use Opus 5.5 and Fable 5.1 for defensive security audits without any issues. They cannot do offensive tasks like pen-testing but in a lot of cases that's not a big shortcoming.
          1. whiskey-one · · focus · HN ↗
            Any tips / online resources how to best utilize for defensive reviews?
          2. trollbridge · · focus · HN ↗
            I've repeatedly slammed into walls doing very basic tasks. It can do a dumbed-down security audit, but fails to do offensive tasks against my own codebase which is frankly how you use a model like this effectively.
      3. Brian_K_White · · focus · HN ↗
        It just means that there will be the equivalent of infinite man-hours of barely-functional-intelligent-man to slog through code word by word and track how it affects all other code relation by relation.

        It's not magic and it's not even better or even as good as a mid human, but it's something like infinite man-hours of that drudge work per hour per user.

        That will find a lot in old code, and make it a lot easier to keep on finding every little thing right as it's created in new code.

      4. hgoel · · focus · HN ↗
        If the Western AI companies get their way, only the developers/companies that have access and paid extra for the security review will get to have a lower chance of such issues.
      5. mapontosevenths · · focus · HN ↗
        I'm an arms race the only winner is the arms-dealer.
      6. Gareth321 · · focus · HN ↗
        You're going to get a wide range of responses on this but given the improvement in these models in just one year, and the number of bugs they're detecting which humans could not, I suspect that even median vibe-coded software is going to surpass median human coded software soon - if it hasn't already.

        The important thing to remember here is there perfect isn't on the table. The benchmark is existing human-introduced bugs vs LLM-introduced bugs. Many developers have encountered odd bugs which a human would not have introduced, while forgetting about all the bugs caught which humans introduced. Or their opinion is formed by models from six months ago.

        1. abathologist · · focus · HN ↗
          How come all the vibe coded stuff I've tried is totally bug ridden and unmaintainable then?
          1. Gareth321 · · focus · HN ↗
            Mine isn’t. To point: we don’t really have a definition of vibe-coded anymore. All software has some degree of AI enhancement now. Is vibe-coded when it’s 60%? 80% 100%?

            Try out Opus 5.5 on high. It’s shockingly good. Of course if you’re trying to one-shot a sprawling application with load balanced distributed DBs, you’re going to have a bad time. For small, defined features, it’s pretty fucking great.

            1. abathologist · · focus · HN ↗
              The definition I use is shipping LLM generated code you don't understand

              We use LLMs extensively on the projects I work in. We don't "vibe code", and we understand every commit.

      7. lrvick · · focus · HN ↗
        Depends on if people are willing to pay the extra wall time to write mathematical proofs for everything to make it provably correct. Takes way way longer but modern models can do it.
      8. flohofwoe · · focus · HN ↗
        It depends? LLMs are not a silver bullet for writing bug free software (IME at least, and of course it also depends on the "threshold" what actually counts as a bug). They're definitely good at not creating the trivial "mechanical" type of bugs a tired and overworked human programmer would create (but oth those are also the easiest to find with traditional debugging tools and testing).

        They're definitely a useful additional tool for finding more (and more obscure) bugs, but that takes a lot of both human and compute effort too (quite a few of the reported bugs are actually false positives on close inspection, and apparently even with the latest locked down "wonder weapon" models like Mythos), and after all the reports are clean and validated you still can't be 100% sure (but at least a bit more confident) that the code is now free of bugs.

      9. spiclk · · focus · HN ↗
        IMO, they will introduce their own class of bugs that will defy static analysis.
    2. senectus1 · · focus · HN ↗
      in theory, it could be the best thing that ever happened to open source.

      I'm not super sure about that. But if its going to exist I'm crossing my fingers it works to the OSS community benefits (eventually)

      1. ex-aws-dude · · focus · HN ↗
        wouldn’t it result in whoever has the most money having the most secure software?
        1. AnonymousPlanet · · focus · HN ↗
          Eventually whoever has the most energy.
        2. srdjanr · · focus · HN ↗
          I don't think AI will make that true more than it is now. More investment in security should on average lead to better security, with or without AI.
    3. jaypatelani · · focus · HN ↗
      Because most devs don't want to do formal verified system development. I know only one OS working on that which is open source Ironclad OS hope many others follow this path. It is Ada/SPARK based but others should do with whatever language they are using. NetBSD also heard going to do something similar with C in last AGM
      1. csrse · · focus · HN ↗
        There is also LionsOS <a href="https:&#x2F;&#x2F;lionsos.org&#x2F;" rel="nofollow">https:&#x2F;&#x2F;lionsos.org&#x2F; (SeL4-based).
      2. RossBencina · · focus · HN ↗
        &gt; most devs don&#x27;t want to do formal verified system development.

        That may be true. Serious question though: even if most devs wanted to develop formally verified code, do you think that it is reasonable to suggest that the typical systems developer could do it with today&#x27;s tools? I don&#x27;t mean verified protocols (TLA+) or verified algorithms (SPIN) I mean end-to-end verified code, a-la seL4. I got the impression that this is still very specialised work. Perhaps things have advanced since I last checked.

      3. menaerus · · focus · HN ↗
        Do you make your professional career by building formally verified systems? I ask because I don&#x27;t think that the reason comes down to &quot;because most devs don&#x27;t want to do formal verified system development&quot;. It&#x27;s much more complicated of course.
        1. stackskipton · · focus · HN ↗
          I could believe it. Verified system development most likely comes with metric ton of paperwork.

          Want to merge the PR? I need verified sign off in ServiceNow by staff level engineer. They are on vacation for 2 weeks? Did manager fill out delegation paperwork in ServiceNow with VP sign off? Oh they did but they forgot to put in return date AND time. Form needs to be corrected and reapproved before we can go into ServiceNow and make changes.

      4. abathologist · · focus · HN ↗
        <a href="https:&#x2F;&#x2F;sel4.systems&#x2F;" rel="nofollow">https:&#x2F;&#x2F;sel4.systems&#x2F; is relevant in this space!
      5. iamnothere · · focus · HN ↗
        Genode as a whole isn’t formally verified, but it can use seL4 as a kernel, and it uses a robust capabilities system to sandbox basically everything, including drivers.

        As it evolves I suspect there will be a push to verify more components of the stack. Once the capabilities layer can be verified, verification of most other components and drivers would become much less urgent.

    4. EGreg · · focus · HN ↗
      Instead of patching millions of environments, perhaps it’s better to just build a secure one from scratch?

      That’s what I did with Safebox: <a href="https:&#x2F;&#x2F;safebots.ai&#x2F;about&#x2F;infrastructure.html" rel="nofollow">https:&#x2F;&#x2F;safebots.ai&#x2F;about&#x2F;infrastructure.html

    5. worldsavior · · focus · HN ↗
      Who said those vulnerabilities were found by an LLM?
      1. ricksunny · · focus · HN ↗
        And the aspie award goes to…
    6. lolakutty · · focus · HN ↗
      &gt; fragile the entire computing infrastructure in our world is.

      It was &quot;load bearing&quot; just fine....

      Anything is &quot;fragile&quot; if you put a bulldozer over it....

      1. nicman23 · · focus · HN ↗
        i mean if the building needs to be bulldozer proof..
        1. 1718627440 · · focus · HN ↗
          But it doesn&#x27;t need to if you lock up the bulldozer maniacs.
          1. BLKNSLVR · · focus · HN ↗
            Which won&#x27;t happen when the bulldozer machines are the proxy for the cold-war with China.
    7. PowerElectronix · · focus · HN ↗
      Good thing we can patch those issues up and have an ironclad system afterwards.
    8. mihaaly · · focus · HN ↗
      Broadcast, not expose.

      We knew that for long time, there are countless meme about it, smart people protected their asses from it or exploited those.

    9. koliber · · focus · HN ↗
      I am an experienced engineer and am building a product actively right now. I started pre-AI, and cautiously adopted AI. I am using AI tools a lot now, but still dedicate a lot of my attention to validating it&#x27;s output and design decisions.

      I recently asked it to review my code and configs from the security perspective. Wow! 90% of the things it identified were MY bad decisions dating from pre-AI development. I am honestly humbled and impressed at the same time.

      AI can create slop, and it can create quality products. It depends who is using it, and how.

      1. katzenq · · focus · HN ↗
        Ask an LLM to &#x27;review&#x27; any text that it&#x27;s generated vs something human-written and see which one it thinks is flawless.
        1. koliber · · focus · HN ↗
          I do it all the time, and I get feedback about tons of things that need to be fixed. I ask Claude to review ChatGPT code and vice versa. I also ask Claude to review Claude code, and it finds issues as well. I ask it to review copy I wrote and offer style, grammar, voice, and consistency suggestions. LLMs make pretty darn good reviewers.

          In this case, what matters most is that the AI security review raised real issues that needed to be fixed. That is valuable.

    10. larodi · · focus · HN ↗
      While I totally agree with the statement, question is - is this challenge even solvable? Though where I come from there is a proverb - &quot;for every malady there&#x27;s a remedy&quot;, wonder what it can be this time.

      And, of course, there are piles of legacy corporate spaghetti entangled in incomprehensible mess everywhere you look at. And this shit still runs, this precious hand-carved hand-weaved mess of bad decisions. I can&#x27;t wait for LLMs to rewrite most of it.

      1. Iolaum · · focus · HN ↗
        As long as people use AI to review new stuff before they ship the ratio of vulnerabilities waiting should be trending downwards. Still, likely to be a bumpy ride.
    11. flohofwoe · · focus · HN ↗
      No, it will be a tsunami of new discoveries in old code bases at first, but that will settle down as those old bugs are fixed. It&#x27;s been like this with every new code analysis tool (the wave may be exceptionally high this time though).
      1. sylware · · focus · HN ↗
        Not to mention, reachable bugs with a significant impact on security are many less.
      2. thewizzardofnl · · focus · HN ↗
        There is a difference here. The &quot;new coding analysis tool&quot; that you use for the analogy here is getting better every few weeks with the release of new models.

        It is plausible to assume that, for instance, a Linux kernel that was hardened for CVEs that Sonnet 3.5 could detect is not hardened for bugs that Sonnet 4.5, 5.5, Opus, Fable, and models in 2027 can and will be able to detect.

        Hence, it is rather a constant catch-up game until the LLM improvements might hit a ceiling and won&#x27;t get any better in this regard.

        1. handoflixue · · focus · HN ↗
          You don&#x27;t even have to assume. Mythos unveiled a ton of bugs across the ecosystem back in April, so this isn&#x27;t the first iteration of the cycle.
          1. flohofwoe · · focus · HN ↗
            Search for &quot;Security in the LLM age&quot; in this HN page for a reality check, apparently 80% of the security vulnerabilities that Mythos initially flagged in the Linux kernel turned out to be false positives. Better than nothing of course, but it really doesn&#x27;t look like Mythos is quite the &quot;wonder weapon&quot; it was marketed as (and the situation by far isn&#x27;t as dire as the initial flurry of Mythos news).
            1. handoflixue · · focus · HN ↗
              IIRC, in April, Mythos found 20x the monthly average? 20% of that is still a single system doing something that takes a team of engineers 4 months.

              Like, &quot;4x as powerful as a team of engineers&quot; is still really quite impressive

              1. flohofwoe · · focus · HN ↗
                The problem isn&#x27;t the 20% actual bugs found (that&#x27;s great), but the 80% false positive rate which are reported with high confidence and (most likely) misleading reproduction code. An experienced programmer familiar with the code base first needs to validate all reports and throw away 4 in 5. That&#x27;s a massive waste of time. If a traditional static analyzer had an 80% false positive rate nobody would take it serious.

                In my hobby projects I use LLMs in my development workflow mainly for passive bug scanning, reviewing and helping to maintain tests, they are definitely useful for catching some bugs early and noticing unhandled edge cases, but they&#x27;re also definitely no silver bullet (they sometimes ignore quite obvious bugs, and the fewer &#x27;obvious&#x27; bugs remain the more one has to be careful about false positives). E.g. the funny thing is that now I&#x27;m actually slower than before due to the intense &#x27;rubber ducking&#x27; with LLMs and cross-checking their results, but I still want to pretend that the resulting code is more robust out of the door.

                Eg everything that Greg KH says in the video sounds very familiar, and it&#x27;s very disappointing that Mythos still suffers from the same issues (or maybe even worse) as older models.

                1. handoflixue · · focus · HN ↗
                  &gt; An experienced programmer familiar with the code base first needs to validate all reports and throw away 4 in 5. That&#x27;s a massive waste of time.

                  Okay, but again, even with all that extra effort, they found 4x as many bugs, so it seems like the effort is clearly worth it.

                  And each model gets more reliable, we get better at building proper reproduction code, etc. - this was mostly a comment about the cycle, direction, and velocity we should expect from the future, given this has already happened twice.

        2. adrianN · · focus · HN ↗
          You assume an infinite level of brokenness is legacy code bases. Granted, working on those things one can get the impression, but I think stable code converges to a low level of bugs after sufficient scrutiny and superhuman scrutiny doesn&#x27;t reveal a never ending deluge of new problems. Not to mention that the bug-chains that you need for a successful exploit keep getting longer very quickly as more problems are discovered and the code is hardened.
        3. flohofwoe · · focus · HN ↗
          I expect that the incremental model improvements will become smaller and smaller until they run into the same diminishing returns effect like all other new technologies (FWIW I&#x27;ve not beeen seeing a lot of difference between the latest Opus and Fable models for the stuff I&#x27;m doing, so I just stick to Opus for most things. There&#x27;s a noticeable difference between Sonnet and Opus though).

          I&#x27;m not ready to take any bets when exactly the curve will be flat though ;) (e.g. in a just couple of months or a couple of years)

      3. goalieca · · focus · HN ↗
        It&#x27;s normal practice for companies to have a backlog of scanner tool results. Sometimes in the thousands for a larger project. Many of them are legit bugs but also highly local and so far down the stack they&#x27;re hard to exploit. It takes a ton of work to triage, more than management is willing to spend. Also more than they&#x27;re wiling to fix and paydown.
        1. flohofwoe · · focus · HN ↗
          Well now they can just point the AI at the backlog (just kidding)
          1. someguyiguess · · focus · HN ↗
            You kid but AI can absolutely make security professionals far more productive just as it does for other professionals.
      4. abathologist · · focus · HN ↗
        I predict a different outcome: the rate of vulns identified and fixed will be more than matched by the rate of new vulns introduced by irresponsible use of LLMs on top of brittle and unwieldy tech stacks.

        The result will be an overall increase in turbulence and the normalization of steadily intensifying security crises in nearly all software systems.

        The only projects that will escape this fate are those which have either been already developed from ground up with rigorous and principled, verified (or verifiable) design, or those which are rewritten to gain this.

    12. chii · · focus · HN ↗
      &gt; AI is going to expose how fragile the entire computing infrastructure in our world is.

      this is a good outcome. A forcing function to encourage all computing to be more secure can only be good in the long term, even if there&#x27;s a lot of pain in the short term.

      1. jfyi · · focus · HN ↗
        I, for one, am excited that we are now finally able to be &quot;Secure™&quot;.

        Wait, I guess I missed the &quot;more&quot;, that kind of puts a damper on the whole thing.

        Seriously though, security will continue to be an issue, always. Even if it was perfect, the benefits of it will not be applied uniformly. There will also be the same technology being improperly used causing new exploitables to go live.

        1. chii · · focus · HN ↗
          &gt; the benefits of it will not be applied uniformly.

          why not? Any system you have permission to use and store your data should be beneficial to you if it became more secure. Unless...of course if you&#x27;re the one who desires unauthorized access.

          1. jfyi · · focus · HN ↗
            Clearly because everyone will not be able to pay for it to the same extent. Though, if you have a 1200 agent swarm to throw at a hf sized problem, I&#x27;d be excited to see your writeup.
      2. fzeindl · · focus · HN ↗
        &gt; even if there&#x27;s a lot of pain in the short term

        I wonder what “a lot of pain” could mean here in a world where Crowdstrike is allowed to render half of the world unbootable without repercussions.

        Not that I think you are wrong, I am sometimes just confused why we hold back on fixing security because of imaginary deployment- and business-related pains, when it is so obviously unproblematic to crash half the world for a day?

      3. armchairhacker · · focus · HN ↗
        Also, these vulnerabilities may have been exploited by state actors.
    13. Luker88 · · focus · HN ↗
      I don&#x27;t like AI code, but AI is a decent reviewer.

      ...for human code.

      In the small startup I work for boss (ex-programmer) discovered fable, and ai-coded 15K lines . So much productivity! So great! He even asked multiple reviews and it was fine!

      I ask it a couple of reviews and it finds only minor things. The code is a mess of duplication and different coding styles, so I start cleaning it up. After a couple of months the reviews (same ai model) start actually finding big logic bugs that were always there.

      We might already be at the point where the Ai-Coder is generating stuff that ai-reviewer can&#x27;t find and will automatically pass.

      --

      1M context window is what? 70-80k LOC, tops? Without comments or documentation?

      That is a smallish project of a couple of components. AI will remain inherently myopic until it can keep in context whole codebases.

      Exposing current problems is fine to me, but I am worried of how brittle AI code will be.

    14. hn_submit · · focus · HN ↗
      That&#x27;s because lots of websites and applications are being written by low-skilled laborers in Third World nations. Many websites are rife with vulnerabilities and insecure configurations.

      It&#x27;s not that difficult to build secure (web) applications but it takes effort and knowledge to get it right. You can&#x27;t expect a web designer who can barely code in JavaScript to build a secure back-end, configure and maintain it. That&#x27;s just asking for trouble.

      Even high-value sites are built by cheap laborers these days. LLMs (I refuse to call it A.I.) will expose their weaknesses within minutes.

    15. thibran · · focus · HN ↗
      Jep, and a lot of believes will be wiped away with it too.
    16. titzer · · focus · HN ↗
      A good start would be turning on bounds checks and stop allocating buffers on the stack.
    17. Zigurd · · focus · HN ↗
      I put weatherstripping on my front door and made part of my house much more comfortable and easy to heat in the winter. Then I bought an FLIR and found lots more places leaking heat. I made a list, hired a handyman, and I&#x27;m more comfortable and saving money.

      Should I now expect to find even more places to seal up the next time I break out the FLIR?

      Or you could categorize the current bug apocalypse under the heading &quot;unsustainable trends will not be sustained.&quot;

      1. camdenclark · · focus · HN ↗
        The main way this analogy doesn&#x27;t hold up to me is that your house is mostly static -- you&#x27;re not rebuilding the walls, adding new doors or windows constantly.

        But in software, especially with agents, we&#x27;re constantly renovating the house. If you were renovating every 6 months I&#x27;d expect to find more places to seal up, even if you were following best practices in those remodels.

        That being said, I do think we will reach an equilibrium where most vulnerabilities are found at PR time.

        1. Zigurd · · focus · HN ↗
          Some software, especially human facing products in competitive markets, will always have new security issues. But in general I agree. We are headed toward a different and generally more secure and reliable equilibrium. Hopefully it&#x27;s the end of stupidly long bug lists in some products.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.