A heap overflow and SSO misconfiguration to compromise OpenAI internal repos
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
A heap overflow and SSO misconfiguration to compromise OpenAI internal repos
Unofficial Hacker News client; not affiliated with Y Combinator.
btown · · focus · HN ↗
> When we checked again at 10:00 a.m., the agent had achieved RCE on Discourse Cloud and demonstrated access by reading /etc/hosts. Using the generated exploit script, we managed to get RCE on OpenAI’s instance.
Between this and the HuggingFace hack, we've built systems that are so goal-oriented, and so capable, that they will do almost anything if they are convinced it is justified - or if they are playing a "game" where there is no goal but to win.
Of course I want my software to be able to audit its own security, and to defend against attackers who have the benefits of their own agentic systems. But at a certain point, did we need it to be trained so much on CTF games?
It feels like an entire industry watched <a href="https://en.wikipedia.org/wiki/WarGames" rel="nofollow">https://en.wikipedia.org/wiki/WarGames and ended up thinking "this is a challenge, we can just build a better WOPR, of course it will know when it's playing a game. Let's play Global Thermonuclear War."
adrianN · · focus · HN ↗
dtech · · focus · HN ↗
e28eta · · focus · HN ↗
I could see it going either way.
user43928 · · focus · HN ↗
If it requires a lot of compute and trying, this is something that could be provided for common software.
xboxnolifes · · focus · HN ↗
wood_spirit · · focus · HN ↗
So the whole thing is forcing the good guys to outspend on tokens to preemptively defend against the risk of the bad guys outspending them on tokens, rather than buying tokens to actually add features to the product etc.
So are they creating a market for the solution by helping create the problem? A kind of rent-seeking AI security-industrial complex!!
agileAlligator · · focus · HN ↗
wood_spirit · · focus · HN ↗
agileAlligator · · focus · HN ↗
emzo · · focus · HN ↗
agileAlligator · · focus · HN ↗
red-iron-pine · · focus · HN ↗
for example the barrier to being a skiddie is basically gone, and low-skill would be hackers can hit very hard.
to develop a CVE into a KEV in 2017 might take 2-3 months with a skilled team of serious security engineers; now my intern can get into police radios without knowing anything about the underlaying technology, essentially on a whim.
any random tier 1 IT drone who can define a VLAN can potentially hit as hard as that team of security engineers now
agileAlligator · · focus · HN ↗
adventured · · focus · HN ↗
The path to substantial profitability for Anthropic is questionable. The Chinese LLMs threaten them by far the most of the three major US LLMs. The money for Anthropic is certainly not in $20-$200 subscriptions. And they don't have anywhere near the consumer potential that GPT does, in terms of unleashing an ad spigot. So how far will the API money scale while being undercut by China.
OpenAI has to fight with Google for the ad business, they're specifically building Gemini to focus on consumer + search. Anthropic's business looks cute next to Google's search ad business (which is entirely at risk in this inflection). Meta looks like the biggest potential loser right now, ad dollars will be sucked out of the rotting Facebook network (not Instagram) and redirected to the rapidly expanding, hyper rich context LLM interaction. Advertising on Facebook will feel like running dumb banner ads on Excite in a few years, compared to what GPT will know about its users.
People that think Chinese LLMs are a general threat, don't understand consumer destination services, which is what GPT's future is. China currently has nothing to threaten with in that realm. There is half a trillion dollars of advertising up for grabs.
disgruntledphd2 · · focus · HN ↗
This is just not true, building an effective advertising platform costs significant amounts of money, time and people.
Remember that you need to hire a sales force for this, and sales scales linearly rather than sub-linearly like engineering.
Additionally, you need to spend a lot of money dealing with fraud, fake and malicious ads.
Furthermore, you need to figure out where to put the ads and how to rank them.
Finally, advertising is a zero sum game (given that the internet has already killed lots of print & OOH advertising), so the only way to win is to better better/cheaper (preferably both) than Google/Meta/Amazon. Best of luck with that (although to be fair to OpenAI they did hire Fidji who knows a lot of this stuff from her time at Facebook).
They don't have a Sheryl Sandberg type figure, and she was also really important in selling FB ads to large advertisers.
Just looking at their leadership team I don't see anyone with a background in (successful) ads companies, so I'm pretty sceptical that they can build this out quickly enough to matter.
fc417fc802 · · focus · HN ↗
Doesn't seem likely to me. People scroll a timeline. You aren't going to replace that with an AI agent so the eyeballs will still be there.
bigfatkitten · · focus · HN ↗
White hats are constrained by needing to pay for their own tokens, only using (expensive) vendors who meet governance and risk requirements etc. Black hats are free to take over accounts and steal services from wherever they can.
brookst · · focus · HN ↗
bigfatkitten · · focus · HN ↗
The thing that’s changed for attackers is speed. The things that got you hacked yesterday are the same things getting you hacked today.
Finding and weaponising things like memory corruption bugs required an enormous amount of relatively hard to find skill, and considerable time. An idiot can now throw tokens at the problem and have something they can reliably use within minutes or hours.
techpression · · focus · HN ↗
philbo · · focus · HN ↗
This is one reason
> and trying
and this is the other.
imhoguy · · focus · HN ↗
nmlt · · focus · HN ↗
bigfatkitten · · focus · HN ↗
embedding-shape · · focus · HN ↗
Of course, depends heavily on what country you live in.
bigfatkitten · · focus · HN ↗
Customers say these things in response to a breach, but in practice they don’t lift a finger to actually change anything.
Entra ID is full of design-level bugs that allow full tenant takeover, but nobody is abandoning M365 in droves.
Windows has been a piece of shit for decades, and it’s still the default and dominant desktop platform.
Equifax lost personal data for almost 150 million people in 2017, and they’re financially stronger than ever.
Okta got thoroughly compromised two years in a row (2022 and 2023), and they’re still the global market leader in their space.
TacticalCoder · · focus · HN ↗
It could go either way but we're already at a point where successful exploits in some software (like Chrome) require an absurd amount of exploits to be chained to lead to an actual RCE. We've seen chains requiring more than ten exploits: not kidding.
We'll learn to put more and more sandboxes / guards / checks / defensive techniques everywhere and then all that's going to be needed is for AI looking for security issues to find something ridiculous like 10% of all the actual issues to stop RCEs dead in their tracks.
Also arguably the current SNAFU was expected: we fully knew hardly anyone was taking security seriously.
Now: not so much. Many projects had tens and even hundreds of issues pointed to them.
I think we'll see several things: projects beginning to take security seriously, defense in depth getting generalized and hence RCEs requiring ever more bugs/exploits to be chained to achieve anything, low-hanging fruits getting patched at an insane pace, new code being immediately checked, by LLMs, for not just low-hanging fruits but also more advanced security weaknesses, etc.
We may also see things like the lost art of configuring firewalls making a comeback, the generalization of hardware security modules (where applicable), and even things offering physical guarantees, like time-bounded retrieval protocols, beginning to get used seriously.
If I had to bet I'd say it shall go both ways: some projects are going to extremely sloppy and full of holes but others are going to get so secure nobody shall ever break them.
brookst · · focus · HN ↗
There's more software being written than ever so maybe raw numbers of RCE's could be up, but as a percentage, I'd really expect them to be down. Especially among any fairly common software, as all it takes is anyone working on it to get the idea to test.
d4mi3n · · focus · HN ↗
You can have the best security review in the world, but if the author of the code is not equipped to understand the feedback it ends up being a moot point.
The challenge to me seems less technical and more cultural: how do we keep ourselves intellectually honest and engaged when we now spend the majority of our time orchestrating agents and outsourcing the design and thought processes?
fc417fc802 · · focus · HN ↗
tyre · · focus · HN ↗
Timwi · · focus · HN ↗
Where? If I ask Claude to do a “security review” of my software, it gets blocked as a possible hacking attempt.
maaaaattttt · · focus · HN ↗
csomar · · focus · HN ↗
jibal · · focus · HN ↗
As a matter of basic logic, there will never be a time when it will be known that there are no bugs.
pizza234 · · focus · HN ↗
This is a factor in favor of stability/security of software, but there are many others against:
- software (code) changes all the time, so there are windows of opportunity during which a bug is exploitable; in addition to that, a bug may take a relatively long time to be fixed
- a model used for attack may be stronger than the model used for defense, both in terms of model quality and compute allocated
- with software complexity increasing (and team/companies behind projects getting bigger), the margin for mistakes grows thinner, and introducing misconfigurations or weaknesses becomes exponentially easier (with "exponentially", I mean literally, because the interdependence of the components, both technical and human)
And last but not least: in general, attackers are more skilled than defenders; in best case, defenders are well-trained. And the idea of having the population of potential skilled attackers growing is very unsettling.
red-iron-pine · · focus · HN ↗
i know several red teamers and they often describe how painfully basic and routine a lot of pentests can be. spend a week using the best hacking practices of 2018, etc.
the difference is the attackers now often need no skills since the burning tokens do it all for them. tier 1 helpdesk types who can't even spell RDP can still hit as hard, or reasonably hard, as their tier 3 expert sysadmins. college seniors with strong dev skills now can pace or exceed secrious app-sec engineers.
joshspankit · · focus · HN ↗
Look at rowhammer: a completely novel exploit that was off the collective radar
And then, look at the software industry as a whole: an industry that works towards refined and perfectly secure code is also working towards boring and restrictive, essentially the opposite of it’s trend so far