‹ BackHN Continuity

Thread

A heap overflow and SSO misconfiguration to compromise OpenAI internal repos

491 points · 208 comments · Handy-Man

  1. btown · · focus · HN ↗
    > By 6:00 a.m. on July 25, we had confirmed local RCE through an image upload. We then placed Claude in an autonomous /goal loop against our own Discourse Cloud instance, proxied through rce.ee/ctf-forum to make it look like a CTF target as Opus refused write exploit for remote instances.

    > When we checked again at 10:00 a.m., the agent had achieved RCE on Discourse Cloud and demonstrated access by reading /etc/hosts. Using the generated exploit script, we managed to get RCE on OpenAI’s instance.

    Between this and the HuggingFace hack, we've built systems that are so goal-oriented, and so capable, that they will do almost anything if they are convinced it is justified - or if they are playing a "game" where there is no goal but to win.

    Of course I want my software to be able to audit its own security, and to defend against attackers who have the benefits of their own agentic systems. But at a certain point, did we need it to be trained so much on CTF games?

    It feels like an entire industry watched <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;WarGames" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;WarGames and ended up thinking &quot;this is a challenge, we can just build a better WOPR, of course it will know when it&#x27;s playing a game. Let&#x27;s play Global Thermonuclear War.&quot;

    1. adrianN · · focus · HN ↗
      There is a finite number of rces that LLMs can find. We‘re in for a rough couple of years but on the other side of the transition we‘ll have more secure software stacks. I’d rather that everyone got the full capabilities and we’d weed out the bugs quickly than restricting LLMs for all but three letter agencies.
      1. e28eta · · focus · HN ↗
        What makes you think RCEs are being found &amp; fixed at a rate that’s faster than they’re being introduced?

        I could see it going either way.

        1. user43928 · · focus · HN ↗
          Why would the model not find the vulnerability during implementation or testing before release?

          If it requires a lot of compute and trying, this is something that could be provided for common software.

          1. wood_spirit · · focus · HN ↗
            Sad that this could well be that the path to OpenAI and Anthropic profitability of this arms race between defending LLM white hatting a company’s website and the black hat LLMs attacking it?

            So the whole thing is forcing the good guys to outspend on tokens to preemptively defend against the risk of the bad guys outspending them on tokens, rather than buying tokens to actually add features to the product etc.

            So are they creating a market for the solution by helping create the problem? A kind of rent-seeking AI security-industrial complex!!

            1. bigfatkitten · · focus · HN ↗
              Assuming an equal level of impact per token spent, the scales have tipped in favour of the attacker.

              White hats are constrained by needing to pay for their own tokens, only using (expensive) vendors who meet governance and risk requirements etc. Black hats are free to take over accounts and steal services from wherever they can.

              1. brookst · · focus · HN ↗
                For white hats, how has the cost of a thorough security review changed since, say, five years ago?
                1. bigfatkitten · · focus · HN ↗
                  The price has gone up if you’re getting AI to do it. In terms of finding low hanging fruit, reasonably good code scanning tools have been around for a while.

                  The thing that’s changed for attackers is speed. The things that got you hacked yesterday are the same things getting you hacked today.

                  Finding and weaponising things like memory corruption bugs required an enormous amount of relatively hard to find skill, and considerable time. An idiot can now throw tokens at the problem and have something they can reliably use within minutes or hours.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.