‹ BackHN Continuity

Thread

A heap overflow and SSO misconfiguration to compromise OpenAI internal repos

491 points · 208 comments · Handy-Man

  1. btown · · focus · HN ↗
    > By 6:00 a.m. on July 25, we had confirmed local RCE through an image upload. We then placed Claude in an autonomous /goal loop against our own Discourse Cloud instance, proxied through rce.ee/ctf-forum to make it look like a CTF target as Opus refused write exploit for remote instances.

    > When we checked again at 10:00 a.m., the agent had achieved RCE on Discourse Cloud and demonstrated access by reading /etc/hosts. Using the generated exploit script, we managed to get RCE on OpenAI’s instance.

    Between this and the HuggingFace hack, we've built systems that are so goal-oriented, and so capable, that they will do almost anything if they are convinced it is justified - or if they are playing a "game" where there is no goal but to win.

    Of course I want my software to be able to audit its own security, and to defend against attackers who have the benefits of their own agentic systems. But at a certain point, did we need it to be trained so much on CTF games?

    It feels like an entire industry watched <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;WarGames" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;WarGames and ended up thinking &quot;this is a challenge, we can just build a better WOPR, of course it will know when it&#x27;s playing a game. Let&#x27;s play Global Thermonuclear War.&quot;

    1. adrianN · · focus · HN ↗
      There is a finite number of rces that LLMs can find. We‘re in for a rough couple of years but on the other side of the transition we‘ll have more secure software stacks. I’d rather that everyone got the full capabilities and we’d weed out the bugs quickly than restricting LLMs for all but three letter agencies.
      1. pizza234 · · focus · HN ↗
        &gt; There is a finite number of rces that LLMs can find.

        This is a factor in favor of stability&#x2F;security of software, but there are many others against:

        - software (code) changes all the time, so there are windows of opportunity during which a bug is exploitable; in addition to that, a bug may take a relatively long time to be fixed

        - a model used for attack may be stronger than the model used for defense, both in terms of model quality and compute allocated

        - with software complexity increasing (and team&#x2F;companies behind projects getting bigger), the margin for mistakes grows thinner, and introducing misconfigurations or weaknesses becomes exponentially easier (with &quot;exponentially&quot;, I mean literally, because the interdependence of the components, both technical and human)

        And last but not least: in general, attackers are more skilled than defenders; in best case, defenders are well-trained. And the idea of having the population of potential skilled attackers growing is very unsettling.

        1. red-iron-pine · · focus · HN ↗
          im not sure i&#x27;d say the attackers are more skilled -- you can get pretty far with the right attitude and a VM running kali linux.

          i know several red teamers and they often describe how painfully basic and routine a lot of pentests can be. spend a week using the best hacking practices of 2018, etc.

          the difference is the attackers now often need no skills since the burning tokens do it all for them. tier 1 helpdesk types who can&#x27;t even spell RDP can still hit as hard, or reasonably hard, as their tier 3 expert sysadmins. college seniors with strong dev skills now can pace or exceed secrious app-sec engineers.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.