‹ BackHN Continuity

Thread

A heap overflow and SSO misconfiguration to compromise OpenAI internal repos

491 points · 208 comments · Handy-Man

  1. btown · · focus · HN ↗
    > By 6:00 a.m. on July 25, we had confirmed local RCE through an image upload. We then placed Claude in an autonomous /goal loop against our own Discourse Cloud instance, proxied through rce.ee/ctf-forum to make it look like a CTF target as Opus refused write exploit for remote instances.

    > When we checked again at 10:00 a.m., the agent had achieved RCE on Discourse Cloud and demonstrated access by reading /etc/hosts. Using the generated exploit script, we managed to get RCE on OpenAI’s instance.

    Between this and the HuggingFace hack, we've built systems that are so goal-oriented, and so capable, that they will do almost anything if they are convinced it is justified - or if they are playing a "game" where there is no goal but to win.

    Of course I want my software to be able to audit its own security, and to defend against attackers who have the benefits of their own agentic systems. But at a certain point, did we need it to be trained so much on CTF games?

    It feels like an entire industry watched <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;WarGames" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;WarGames and ended up thinking &quot;this is a challenge, we can just build a better WOPR, of course it will know when it&#x27;s playing a game. Let&#x27;s play Global Thermonuclear War.&quot;

    1. adrianN · · focus · HN ↗
      There is a finite number of rces that LLMs can find. We‘re in for a rough couple of years but on the other side of the transition we‘ll have more secure software stacks. I’d rather that everyone got the full capabilities and we’d weed out the bugs quickly than restricting LLMs for all but three letter agencies.
      1. jibal · · focus · HN ↗
        No one with a shred of intellectual integrity uses a &quot;There is a finite number&quot; strawman.

        As a matter of basic logic, there will never be a time when it will be known that there are no bugs.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.