A heap overflow and SSO misconfiguration to compromise OpenAI internal repos
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
A heap overflow and SSO misconfiguration to compromise OpenAI internal repos
Unofficial Hacker News client; not affiliated with Y Combinator.
btown · · focus · HN ↗
> When we checked again at 10:00 a.m., the agent had achieved RCE on Discourse Cloud and demonstrated access by reading /etc/hosts. Using the generated exploit script, we managed to get RCE on OpenAI’s instance.
Between this and the HuggingFace hack, we've built systems that are so goal-oriented, and so capable, that they will do almost anything if they are convinced it is justified - or if they are playing a "game" where there is no goal but to win.
Of course I want my software to be able to audit its own security, and to defend against attackers who have the benefits of their own agentic systems. But at a certain point, did we need it to be trained so much on CTF games?
It feels like an entire industry watched <a href="https://en.wikipedia.org/wiki/WarGames" rel="nofollow">https://en.wikipedia.org/wiki/WarGames and ended up thinking "this is a challenge, we can just build a better WOPR, of course it will know when it's playing a game. Let's play Global Thermonuclear War."
wood_spirit · · focus · HN ↗
I’m in the “glorified spell checker” camp, although I don’t mean to reduce their impressive utility and belittle them in the way many people read that term and infer.
So I am not sure that an llm “justifies” anything. I mean that their “thinking” text talks about justifications but it is just a very advanced statistical regurgitation of the kind of text humans use. I don’t think it means the model has internalised the meaning of it (as witness when you talk to an llm how often it forgets what you recently told it was important etc).
What you really have is a model that tries the statistically most probable thing to say next and so on and what is really cool is how effective this is at generating a path that we can slap a narrative over afterwards that makes the whole thing feel motivated and consistent, like the model started off knowing how it was going to get to the destination.
Which is, under the hood, a completely different kind of “intelligence” as the supercomputer in War Games.
Certhas · · focus · HN ↗
Likewise, the fact that LLMs are a stochastic autoregressive process (which is a class of systems every bit as rich as the ODEs used to model neurons) tells us nothing a priori.
wood_spirit · · focus · HN ↗
Certhas · · focus · HN ↗
If I give an LLM to compact its context window, so the context it carries can evolve iteratively over time as more and more things come in, is that enough?
Compacting the context is really a very, very interesting example here. The "next token predictor" is telling an external tool to change all "previous" tokens. So an LLM + a harness that allows compacting the context is no longer just a token predictor at all!
You don't need continuous learning to get interesting dynamics. You just need feedback loops.
wood_spirit · · focus · HN ↗
Richard Dawkins says he thinks LLMs think.
And the physical angle is that nothing is special about humans and software simulating it would also be thinking.
But from using LLMs all the time, and understanding what is under the hood, I’m thinking that the current approaches aren’t cutting it for me and I’m not expecting them to get there. There was such big jumps early on but progress is slowing as though diminishing returns.
So we can build things that think and outthink us, but I don’t think anything we’ve hit upon yet is going to scale up into it.
leg100 · · focus · HN ↗
They're not comparable.