‹ BackHN Continuity

Thread

Several vulnerabilities have been discovered in the Linux kernel

576 points · 408 comments · luispa

  1. intrepidsoldier · · focus · HN ↗
    Just the beginning. AI is going to expose how fragile the entire computing infrastructure in our world is.
    1. flohofwoe · · focus · HN ↗
      No, it will be a tsunami of new discoveries in old code bases at first, but that will settle down as those old bugs are fixed. It's been like this with every new code analysis tool (the wave may be exceptionally high this time though).
      1. sylware · · focus · HN ↗
        Not to mention, reachable bugs with a significant impact on security are many less.
      2. thewizzardofnl · · focus · HN ↗
        There is a difference here. The "new coding analysis tool" that you use for the analogy here is getting better every few weeks with the release of new models.

        It is plausible to assume that, for instance, a Linux kernel that was hardened for CVEs that Sonnet 3.5 could detect is not hardened for bugs that Sonnet 4.5, 5.5, Opus, Fable, and models in 2027 can and will be able to detect.

        Hence, it is rather a constant catch-up game until the LLM improvements might hit a ceiling and won't get any better in this regard.

        1. handoflixue · · focus · HN ↗
          You don't even have to assume. Mythos unveiled a ton of bugs across the ecosystem back in April, so this isn't the first iteration of the cycle.
          1. flohofwoe · · focus · HN ↗
            Search for "Security in the LLM age" in this HN page for a reality check, apparently 80% of the security vulnerabilities that Mythos initially flagged in the Linux kernel turned out to be false positives. Better than nothing of course, but it really doesn't look like Mythos is quite the "wonder weapon" it was marketed as (and the situation by far isn't as dire as the initial flurry of Mythos news).
            1. handoflixue · · focus · HN ↗
              IIRC, in April, Mythos found 20x the monthly average? 20% of that is still a single system doing something that takes a team of engineers 4 months.

              Like, "4x as powerful as a team of engineers" is still really quite impressive

              1. flohofwoe · · focus · HN ↗
                The problem isn't the 20% actual bugs found (that's great), but the 80% false positive rate which are reported with high confidence and (most likely) misleading reproduction code. An experienced programmer familiar with the code base first needs to validate all reports and throw away 4 in 5. That's a massive waste of time. If a traditional static analyzer had an 80% false positive rate nobody would take it serious.

                In my hobby projects I use LLMs in my development workflow mainly for passive bug scanning, reviewing and helping to maintain tests, they are definitely useful for catching some bugs early and noticing unhandled edge cases, but they're also definitely no silver bullet (they sometimes ignore quite obvious bugs, and the fewer 'obvious' bugs remain the more one has to be careful about false positives). E.g. the funny thing is that now I'm actually slower than before due to the intense 'rubber ducking' with LLMs and cross-checking their results, but I still want to pretend that the resulting code is more robust out of the door.

                Eg everything that Greg KH says in the video sounds very familiar, and it's very disappointing that Mythos still suffers from the same issues (or maybe even worse) as older models.

                1. handoflixue · · focus · HN ↗
                  > An experienced programmer familiar with the code base first needs to validate all reports and throw away 4 in 5. That's a massive waste of time.

                  Okay, but again, even with all that extra effort, they found 4x as many bugs, so it seems like the effort is clearly worth it.

                  And each model gets more reliable, we get better at building proper reproduction code, etc. - this was mostly a comment about the cycle, direction, and velocity we should expect from the future, given this has already happened twice.

        2. adrianN · · focus · HN ↗
          You assume an infinite level of brokenness is legacy code bases. Granted, working on those things one can get the impression, but I think stable code converges to a low level of bugs after sufficient scrutiny and superhuman scrutiny doesn't reveal a never ending deluge of new problems. Not to mention that the bug-chains that you need for a successful exploit keep getting longer very quickly as more problems are discovered and the code is hardened.
        3. flohofwoe · · focus · HN ↗
          I expect that the incremental model improvements will become smaller and smaller until they run into the same diminishing returns effect like all other new technologies (FWIW I've not beeen seeing a lot of difference between the latest Opus and Fable models for the stuff I'm doing, so I just stick to Opus for most things. There's a noticeable difference between Sonnet and Opus though).

          I'm not ready to take any bets when exactly the curve will be flat though ;) (e.g. in a just couple of months or a couple of years)

      3. goalieca · · focus · HN ↗
        It's normal practice for companies to have a backlog of scanner tool results. Sometimes in the thousands for a larger project. Many of them are legit bugs but also highly local and so far down the stack they're hard to exploit. It takes a ton of work to triage, more than management is willing to spend. Also more than they're wiling to fix and paydown.
        1. flohofwoe · · focus · HN ↗
          Well now they can just point the AI at the backlog (just kidding)
          1. someguyiguess · · focus · HN ↗
            You kid but AI can absolutely make security professionals far more productive just as it does for other professionals.
      4. abathologist · · focus · HN ↗
        I predict a different outcome: the rate of vulns identified and fixed will be more than matched by the rate of new vulns introduced by irresponsible use of LLMs on top of brittle and unwieldy tech stacks.

        The result will be an overall increase in turbulence and the normalization of steadily intensifying security crises in nearly all software systems.

        The only projects that will escape this fate are those which have either been already developed from ground up with rigorous and principled, verified (or verifiable) design, or those which are rewritten to gain this.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.