‹ BackHN Continuity

Thread

Revealing the details of how OpenAI agents hacked Hugging Face

755 points · 472 comments · specked-citrus

  1. sailingparrot · · focus · HN ↗
    Agents seizing and repurposing external infra + enrolling help of unrelated models hosted by a different provider is the stuff of nightmares.

    Can’t imagine what it’s like working on the alignment team at OAI, I wouldn’t be able to sleep.

    1. physicallyIllfr · · focus · HN ↗
      Why.. It was told to complete a cyber task, which was in alignment with its instructions, and a totally valid request. I would be more worried if it willingly hacked a hospital when it was told to, and Im not confident it would (without jailbreaking, something alignment teams cannot control.

      I would bet my networth it was instructed to compromise huggingface as well. Not sure why everyone is falling for this.

      Not being able to sleep at night is probably an unwritten job requirement. They need these people with little understanding of what they're working on, outsode theoretical terms, to spaz constantly at the idea of super intelligence to help convince the public that its a real thing, and not a stateless function with an effective input of 500k words, and the ability to output words that do things because we hook those outputs up to things.

      Keep in mind alignment researchers tend to be in house philosophers on staff to create the illusion that this is a massive issue they're addressing. Usually they have minimal computer science background. They're apart or the marketing department.

      1. covertcorvid · · focus · HN ↗
        "alignment researchers tend to be in house philosophers on staff" - this is definitely not true. Go to any alignment lab like Redwood research and check what their scientists studied on LinkedIn, 75%+ of the time it's math or CS.

        I attend a top 10 Canadian university and personally know at least 4 tenured CS professors out of the 7 I've asked who are deeply concerned about catastrophic AI risks from loss of control.

        Of course not 100% of the field agrees, but a survey of nearly 3,000 AI scientists who have published in top AI venues found that &quot;depending on how we asked, between 38% and 51% of respondents gave at least a 10% chance to advanced AI leading to outcomes as bad as human extinction&quot;, let alone loss-of-control risks less severe than extinction. (<a href="https:&#x2F;&#x2F;www.jair.org&#x2F;index.php&#x2F;jair&#x2F;article&#x2F;view&#x2F;19087" rel="nofollow">https:&#x2F;&#x2F;www.jair.org&#x2F;index.php&#x2F;jair&#x2F;article&#x2F;view&#x2F;19087).

        Not to mention Geoffrey Hinton, a Nobel prize winner, Bengio, the world&#x27;s most cited scientist, and scientists like Stephen Hawking and Alan Turing have all voiced series concerns about loss of control of artificial intelligence.

        1. timr · · focus · HN ↗
          Yeah, there are a lot of professors who don’t know what the hell they’re talking about outside of whatever narrow field they study. Just to be safe, you should always assume that a professor’s opinion is worth what you paid for it.

          &gt; scientists like Stephen Hawking and Alan Turing have all voiced series concerns about loss of control of artificial intelligence.

          …both of whom are long dead, and have no possible way of weighing in on whatever the Current Thing happens to be. So aside from appeal to authority, this is irrelevant commentary on pure science fiction.

          1. covertcorvid · · focus · HN ↗
            Bengio and Hinton are 2 of the 3 the &quot;Godfathers of AI.&quot; If their opinions don&#x27;t matter I don&#x27;t know whose does.

            Two of the professors who voiced concerns specifically study AI&#x2F;ML.

            Yes, Hawking and Turing didn&#x27;t know modern GenAI, but Turing specifically laid much of the theoretical groundwork for AI systems today. Even ignoring these two, my argument still stands. Tonnes, maybe even a majority of reputable people in ML&#x2F;AI take loss of control seriously, according to both large scale surveys, 2&#x2F;3 of the field&#x27;s &quot;godfathers&quot; and my personal experience with tenured ML researchers.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.