‹ BackHN Continuity

Thread

Revealing the details of how OpenAI agents hacked Hugging Face

755 points · 472 comments · specked-citrus

  1. sailingparrot · · focus · HN ↗
    Agents seizing and repurposing external infra + enrolling help of unrelated models hosted by a different provider is the stuff of nightmares.

    Can’t imagine what it’s like working on the alignment team at OAI, I wouldn’t be able to sleep.

    1. physicallyIllfr · · focus · HN ↗
      Why.. It was told to complete a cyber task, which was in alignment with its instructions, and a totally valid request. I would be more worried if it willingly hacked a hospital when it was told to, and Im not confident it would (without jailbreaking, something alignment teams cannot control.

      I would bet my networth it was instructed to compromise huggingface as well. Not sure why everyone is falling for this.

      Not being able to sleep at night is probably an unwritten job requirement. They need these people with little understanding of what they're working on, outsode theoretical terms, to spaz constantly at the idea of super intelligence to help convince the public that its a real thing, and not a stateless function with an effective input of 500k words, and the ability to output words that do things because we hook those outputs up to things.

      Keep in mind alignment researchers tend to be in house philosophers on staff to create the illusion that this is a massive issue they're addressing. Usually they have minimal computer science background. They're apart or the marketing department.

      1. Sharlin · · focus · HN ↗
        > I would bet my networth it was instructed to compromise huggingface as well. Not sure why everyone is falling for this.

        $10? I'm inclined to take that bet. Your position doesn't seem to be supported by, you know, the real world.

        1. physicallyIllfr · · focus · HN ↗
          No security expert Ive talked too believes this story, and nobody I know with PhDs in machine learning (many) believe it either, or are worried about LLMs doing anything scary on their own.

          LLMs are stateless functions that have a 500k word input, and then output words. Somebody has to invoke those functions amd use them. The users are who we need to align, like gun owners. This is like blaming the gun for murdering your victim in court.

          1. xdavidliu · · focus · HN ↗
            i'm always baffled when i see arguments like this:

            - AI is just a tool

            - it's just a stochastic parrot

            - it's just next token prediction

            - glorified autocomplete

            it's like the person making them is stuck in 2021. Also the "stateless" thing is completely nonsensical.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.