‹ BackHN Continuity

Thread

A heap overflow and SSO misconfiguration to compromise OpenAI internal repos

491 points · 208 comments · Handy-Man

  1. btown · · focus · HN ↗
    > By 6:00 a.m. on July 25, we had confirmed local RCE through an image upload. We then placed Claude in an autonomous /goal loop against our own Discourse Cloud instance, proxied through rce.ee/ctf-forum to make it look like a CTF target as Opus refused write exploit for remote instances.

    > When we checked again at 10:00 a.m., the agent had achieved RCE on Discourse Cloud and demonstrated access by reading /etc/hosts. Using the generated exploit script, we managed to get RCE on OpenAI’s instance.

    Between this and the HuggingFace hack, we've built systems that are so goal-oriented, and so capable, that they will do almost anything if they are convinced it is justified - or if they are playing a "game" where there is no goal but to win.

    Of course I want my software to be able to audit its own security, and to defend against attackers who have the benefits of their own agentic systems. But at a certain point, did we need it to be trained so much on CTF games?

    It feels like an entire industry watched <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;WarGames" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;WarGames and ended up thinking &quot;this is a challenge, we can just build a better WOPR, of course it will know when it&#x27;s playing a game. Let&#x27;s play Global Thermonuclear War.&quot;

    1. wood_spirit · · focus · HN ↗
      &gt; they will do almost anything if they are convinced it is justified

      I’m in the “glorified spell checker” camp, although I don’t mean to reduce their impressive utility and belittle them in the way many people read that term and infer.

      So I am not sure that an llm “justifies” anything. I mean that their “thinking” text talks about justifications but it is just a very advanced statistical regurgitation of the kind of text humans use. I don’t think it means the model has internalised the meaning of it (as witness when you talk to an llm how often it forgets what you recently told it was important etc).

      What you really have is a model that tries the statistically most probable thing to say next and so on and what is really cool is how effective this is at generating a path that we can slap a narrative over afterwards that makes the whole thing feel motivated and consistent, like the model started off knowing how it was going to get to the destination.

      Which is, under the hood, a completely different kind of “intelligence” as the supercomputer in War Games.

      1. Arn_Thor · · focus · HN ↗
        I used to share that perspective until very recently, but today I think it&#x27;s an outdated way to think of the cutting-edge LLMs. There is so much more going on, with MOEs, internal loops, guardrails and tools that I suspect we&#x27;re dealing with something that&#x27;s a little more than the sum of its parts. Not intelligent in the way we recognize in biological organisms, but certainly something beyond a mere Markov chain.
        1. HarlequinHair · · focus · HN ↗
          Make no mistakes.

          LLMs are language model, and nowhere in their code you can find actual reasoning. Re-reinforcement is not magical process that builds conscience or emotions.

          We are talking about probability built on statistics, with extea steps.

          Stop humanizing LLMs.

          1. jibal · · focus · HN ↗
            Agents are not simple language models.

            You can&#x27;t find actual reasoning in a brain either. (Note that you can&#x27;t tell the difference between a conscious brain and a comatose brain by examining them.) This is the same as Leibniz&#x27;s mill argument ... it&#x27;s a fallacy of composition.

            &gt; Re-reinforcement is not magical process that builds conscience or emotions.

            They aren&#x27;t the result of magic at all, but we are nowhere near the point of identifying what processes do or don&#x27;t produce consciousness (or a conscience) or can be characterized as having emotions.

            &gt; Stop humanizing LLMs.

            That&#x27;s a clearly dishonest mischaracterization of the GP.

            I&#x27;ve read some of your other comments about LLMs and I find them unreasonably reductionistic, whereas I think the word &quot;just&quot; should be banned from ontological discussion, so I don&#x27;t think further engagement would be beneficial and I won&#x27;t be engaging in it. (And I&#x27;m actually quite conservative in ascribing cognitive traits to LLMs or other &quot;AI&quot;.)

            1. HarlequinHair · · focus · HN ↗
              The best non technical explanation you can give is &quot;An AI agent is an LLM that can take actions&quot;.

              While an agent doesn&#x27;t necessarily have to be powered by an LLM, most modern AI agents are.

              You pointing at a human brain does not change that an AI agent is not intelligent and cannot think, we are still talking about probability built on statistics with extra steps.

              I am not trying to be dishonest, we should stop making analogies between AI and actual thinking, because they are two entire different concepts.

              Who developed these technologies used the words &quot;thinking&quot; and &quot;reasoning&quot;, this does not mean they are actually thinking and reasoning. Somewhere you still have a processor calculating, with no empathy.

              So, again: stop humanizing AI. This sentence shouldn&#x27;t make you angry.

              1. pizza234 · · focus · HN ↗
                &gt; we are still talking about probability built on statistics with extra steps.

                There is a wrong assumption here: confusing primitives with emergent properties.

                One can&#x27;t look at the primitivies and assume that certain properties will not emerge. It would be exactly like looking at aminoacids and state that intelligence can&#x27;t develop from them.

                &gt; You pointing at a human brain does not change that an AI agent is not intelligent and cannot think

                That depends on the definition of intelligence and thinking, and it is dishonest not to give any definition (and most importantly, one that is not human-centered).

                AIs are currently fulfilling several aspects of intelligence and thinking, by any defition of intelligence. If you don&#x27;t notice that, it&#x27;s just because you have informed yourself enough. Having said that, I don&#x27;t doubt that there are aspects that AI are lacking (e.g. retention&#x2F;plasticity&#x2F;perception), but the line is blurry, and they&#x27;re advancing (too) fast.

                Empathy is actually a very important aspect of the AI problems, but it&#x27;s not part of intelligence. Sociopaths don&#x27;t have it, and yet, you wouldn&#x27;t doubt that they&#x27;re intelligent.

                1. tripzilch · · focus · HN ↗
                  &gt; It would be exactly like looking at aminoacids and state that intelligence can&#x27;t develop from them.

                  You do realize that amino acids exist on a scale some orders of magnitude smaller than the gates we build GPUs out of?

                  Honestly, this &quot;you could say the same about humans&quot;-argument is getting so tired. A brain neuron is so complicated, we can&#x27;t even simulate a single one ...

                  At the very least there is no reason why you should jump to a human brain, of all things.

                  But the whole argument kinda loses its spice, when you say &quot;well you could say the same about a mouse brain&quot;, and you know what happens when you create swarms of 1000s of mice ... super intelligence, right?

                  1. anonymars · · focus · HN ↗
                    Do mice satisfy your definition of intelligent?
                    1. tripzilch · · focus · HN ↗
                      That&#x27;s the point.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.