‹ BackHN Continuity

Thread

Greg Kroah-Hartman – Security in the LLM Age [video]

337 points · 128 comments · usernomdeguerre

  1. usernomdeguerre · · focus · HN ↗
    Greatly appreciated the candor. I've included a few slides into text that i thought were eye-opening to me:

    From his Kernel Recipes 2026 slide on Mythos

    ```

      Mythos's 79 vulnerabilities:
      24 - no detail at all "something crashed"
      14 - not a bug at all
      3 - totally made up data
      15 - already fixed in latest release
        - 11 by others
        - 4 by anthropic
      20 - fixes were needed
        - 7 "assume a malicious filesystem image"
        - 2 "assume you can inject a malicious network packet into the middle of the stack"
        - 2 "NOMMU"
        - 6 sctp networking issues for untrusted devices
        - 2 ipv6 minor network issues 
        - 1 gpu driver for local malicious user
    
    ```

    GHK called this "10 'real' bugfixes", which to me sounds like there's a wild hype machine around these companies and uncritical parroting of every press release they make that falls apart when you engage the affected real experts.

    1. OtherShrezzing · · focus · HN ↗
      We’ve seen this in a few open source repos we voluntarily manage security on. They’re not massive repos, but big enough they get attention from security researchers.

      Most of the real low hanging fruit was picked up by humans years ago. When doing automated scanning, the majority of stuff is overly-verbose nonsense which takes hours of expert human labour to understand, test, and discard.

      Reading through a Claude generated false positive is absolutely excruciating, because it is absolutely determined that what it’s found is justified. Often you’ll receive very long accompanying “proof of concept” code which demonstrates absolutely wild scenarios. It’s especially frustrating when you’re volunteering your time for a project, and a well-meaning contributor submits the report without the technical nous to understand why you’re rejecting it.

      1. b112 · · focus · HN ↗
        Right now, all top tier LLMs are as eager, bright 20ish year old interns.

        Very gung ho, full of energy, loads of book learning, no real world experience or understanding of why things are as they are.

        Leave them to their own devices at your peril. Trust nothing they do.

        Yet directly guide them, monitor everything they do, some value emerges.

        1. charcircuit · · focus · HN ↗
          Have you used a frontier model since 2025? You are underplaying their strength.
          1. fc417fc802 · · focus · HN ↗
            Okay so now they're like a top percentile fresh grad on meth. Still a lack of real world experience plus some bizarre failures that illustrate gaping holes in the mental model. Does that description work for you?
            1. TeMPOraL · · focus · HN ↗
              Well, so basically supercharged fresh grad. CS/math implied.

              In human terms, that's already at least a standard deviation above average person.

          2. 12376 · · focus · HN ↗
            Kroah-Hartmann has used the closed frontier++ model, and it made up 37 out of 76 vulnerabilities.
            1. Gigachad · · focus · HN ↗
              And the bugs that actually were real rely on a setup so contrived it’s unlikely anyone in the world is impacted.

              It’s still good that some real bugs are being patched but what is being reported to the media is so overblown.

          3. crote · · focus · HN ↗
            We've been seeing "But you're not using the latest model!" over and over again, with every new model supposedly "groundbreaking" and "a game-changer" - just for the general population to conclude a few months later that it once again doesn't live up to the crazy marketing hype.

            Anthropic claimed that Mythos was so good at finding vulnerabilities that it was too dangerous to release to the public. As this post clearly shows: that (only again) simply isn't true. If you believe your favorite flavor of frontier model is the exception, it is up to you to provide proof to back up that claim.

        2. verdverm · · focus · HN ↗
          I sure hope they stop training them to be spaghetti throwers, please stop and ask questions when there is ambiguity
          1. spwa4 · · focus · HN ↗
            Well let's see. You sell to management. Does management buy:

            1) nuanced tools that talk back, question assumptions, take over decisions, ... oh and expose just how much management knows about the business. Or how little)

            2) a tool that can provide the excuse "we've had our source checked and dealt with the remarks"

            We all know the answer.

        3. truncate · · focus · HN ↗
          I wish people stopped equating LLM to interns or junior engineers. They are tools, and as good they may be at some specific things people do, they also really suck at many more which we wouldn’t find acceptable in humans.
          1. b112 · · focus · HN ↗
            It's merely a way to frame things, in terms of experience and trust. And it highlights how an LLM can code very well, but not truly understand the ramifications of that code.
            1. t_mahmood · · focus · HN ↗
              An "eager, bright 20ish year old interns" will grow as you guide them, while LLM will not, they'll only grow when their owner (definitely not us) update them. So, trying to humanizing some tool with human emotion is wrong way to framing it. Treat tool as tool.
          2. darkwater · · focus · HN ↗
            Indeed they are tools. But if people/companies treat them as tools that can take a problem a human used to solve, and make them solve it from start to end, then the comparison begins to be necessary.
          3. skinfaxi · · focus · HN ↗
            I wonder if the same was once said about calculating machines and calculators.
        4. delusional · · focus · HN ↗
          > Right now, all top tier LLMs are as eager, bright 20ish year old interns. > Yet directly guide them, monitor everything they do, some value emerges.

          That's simply not true. I've had some very talented interns, and they are leagues ahead of what the LLMs can do. Not that it matters though, because the point of having interns wasn't to have them produce value. What made the investment worth it was that 12 months down the line I would have a competent colleague that I could have an interesting conversation with. A human person that could challenge some of my blind spots. A person that could take responsibility of something. Maybe not my most important work, but some of it. You don't get ANY of that from the LLM.

        5. Zigurd · · focus · HN ↗
          In some ways the reality is worse: the same version of an LLM won't get better at its job, even though you might get better at prompting it. Newer versions are trained on more code, which has obvious benefits, and their harnesses are better at taking advantage of tools that were created to keep human coders out of trouble.

          The other side of the coin is that coding agents are not maximally productive unless you give them enough rope to potentially hang themselves. Over roughly the past year, the coding agents I use have gone from hot garbage to pretty consistently useful, especially if I find tasks where I can give them a lot of running room. On the other hand, last week I found a case where the coding agent was looping and flailing like it was doing every third try a year ago.

          They fail less often, but they fail in the same way.

      2. bitwize · · focus · HN ↗
        Before:

        There is a vulnerability in library X when you call Y with these specially crafted parameters; as seen in the attached logs it can overflow buffer Z and clobber memory potentially leading to an RCE.

        After:

        Honest take: the load-bearing constraint violation is real. The log documenting the exploit gate is the ledger which weaves the story.

        1. prox · · focus · HN ↗
          Blast radius!
          1. bitwize · · focus · HN ↗
            Now you see the whole picture!
      3. IshKebab · · focus · HN ↗
        I think the way to handle this is to just feed it into another AI agent (a better one) and ask it how serious the issue actually is. Fight fire with fire!
        1. Juliate · · focus · HN ↗
          That&#x27;s somehow what <a href="https:&#x2F;&#x2F;sashiko.dev" rel="nofollow">https:&#x2F;&#x2F;sashiko.dev is.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.