‹ BackHN Continuity

Thread

A heap overflow and SSO misconfiguration to compromise OpenAI internal repos

491 points · 208 comments · Handy-Man

  1. btown · · focus · HN ↗
    > By 6:00 a.m. on July 25, we had confirmed local RCE through an image upload. We then placed Claude in an autonomous /goal loop against our own Discourse Cloud instance, proxied through rce.ee/ctf-forum to make it look like a CTF target as Opus refused write exploit for remote instances.

    > When we checked again at 10:00 a.m., the agent had achieved RCE on Discourse Cloud and demonstrated access by reading /etc/hosts. Using the generated exploit script, we managed to get RCE on OpenAI’s instance.

    Between this and the HuggingFace hack, we've built systems that are so goal-oriented, and so capable, that they will do almost anything if they are convinced it is justified - or if they are playing a "game" where there is no goal but to win.

    Of course I want my software to be able to audit its own security, and to defend against attackers who have the benefits of their own agentic systems. But at a certain point, did we need it to be trained so much on CTF games?

    It feels like an entire industry watched <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;WarGames" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;WarGames and ended up thinking &quot;this is a challenge, we can just build a better WOPR, of course it will know when it&#x27;s playing a game. Let&#x27;s play Global Thermonuclear War.&quot;

    1. paimapi · · focus · HN ↗
      &gt;they will do almost anything if they are convinced it is justified - or if they are playing a &quot;game&quot; where there is no goal but to win

      just a small caution on this anthropomorphism - it implies there&#x27;s some high-order &#x27;thinking&#x27; behind it. in reality, it&#x27;s probably healthier to see LLMs as a combination of symbolic logic reasoning steps paired with probabilistic token predictor generating the proponents and operators in that chain, all trained by humans on different large data sets

      to be &#x27;goal-oriented&#x27; implies that there&#x27;s the capacity to be anything else and I don&#x27;t think that&#x27;s how LLMs operate at all. I think they only know how to operate within their design parameters and much of that design is simply much further upstream during the training and post-training processes. that opacity makes it feel like &#x27;intelligence&#x27; when you&#x27;re interacting with it as a downstream product because you&#x27;ll see an agent act in a way that you didn&#x27;t command - but that&#x27;s simply a result of your not being shown all the antecedent mappings and architectural design

      something something indistinguishable from magic as that one guy said

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.