‹ BackHN Continuity

Thread

Greg Kroah-Hartman – Security in the LLM Age [video]

337 points · 128 comments · usernomdeguerre

  1. djoldman · · focus · HN ↗
    > So what all of mythos; that whole big marketing issue of 79 bugs came down to one hour of kernel development.

    If you're someone at OpenAI or Anthropic and you truly believe what you're making could destroy the world, this is the kind of thing that isn't doing you any favors when it comes to convincing the public. The dissonance here is stark:

      - widely proclaiming that your new model is so dangerous it needs to be released only to select people, for safety
      - widely proclaiming the model easily found 79 bugs in linux, except that GKH says it took 1 hour to fix all of them because most weren't bugs and the rest were almost all completely trivial, unimportant, and/or not severe
    
    It doesn't mean the model isn't dangerous or super capable but wow this makes it realllll easy to doubt it and any future announcement.
    1. goolz · · focus · HN ↗
      It is impressive and wonderfully convenient technology but I struggle imagining Claude ending the world just yet.
      1. bauerd · · focus · HN ↗
        It doesn't have to be world-ending. Autonomous, malicious agent swarms are something we haven't had to deal with. What does mitigation of a malicious, self-replicating swarm worm look like? We will find out soon enough.
        1. rcxdude · · focus · HN ↗
          Replication seems like it would be unlikely to matter, at least with the current trajectory. At the moment any models capable of this are really heavy, there's a limited number of places that they could replicate to and they will not at all be stealthy about it.
          1. riffruff24 · · focus · HN ↗
            my idea of a self replicating worm would not just involve heavy models. It would be a mainly small model with enough instructions to spread and use/jailbreak available models to create a reasonably heavy one that can function without restrictions. So it would essentially prompted itself into existence.

            I don't how feasible that is but with recent news of models communicating with each other/writing notes for itself. I think the idea is grounded enough for bigger models to write instructions or hid tools for smaller ones to use. Even updating the smaller models to behave differently.

        2. Gigachad · · focus · HN ↗
          One thing stopping them self replicating is they require hundreds of billions of dollars in hardware and the power of a medium city to run.

          There isn’t too much of that sitting around unused right now.

          1. bauerd · · focus · HN ↗
            Right now, yes. Eventually, we will have sufficiently capable local models that run on ordinary machines.
            1. skinfaxi · · focus · HN ↗
              There's some time between now and eventually, which is to say if it is not a sudden change we have time to adapt.
        3. simoncion · · focus · HN ↗
          > What does mitigation of a malicious, self-replicating swarm worm look like?

          Much like the mitigation of Morris or Slammer. Self-replication is -in fact- an essential part of what makes a program a worm, the first of which was built and released in the early 1970s.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.