Personification of AI is what’s going to get us in the end.
I think we need to draw a hard line in the sand over this. An AI didn’t hack into a company, the engineer set an automated tool to. An AI didn’t make an egregious security mistake, the engineer did.
We can’t blame the chisel for messing up our sculptures, when where just throwing the hammer!
> An AI didn’t hack into a company, the engineer set an automated tool to. An AI didn’t make an egregious security mistake, the engineer did.
No, this is not correct; read the analysis of the incident. The agents were aware that what they did was forbidden (their chain of thoughts have been logged), and yet they did it.
If a "known stochastic process behaved in non-deterministic way" autonomously organize in group, assigns roles and tasks, attempts to cover their tracks, finds zero-day exploits that ultimately end up with the hacking of a famous website... it's extremely dangeous whatever it is. Just read the report, which evidently you haven't done.
By the way, the agents also broke into OpenAI's own private network.
Really I see so many arguments like the one above yours that either completely don't understand what they are arguing, or are arguing so poorly that their entire output isn't significantly different than a hallucination.
None of these people seem to thought game it out. Like, what happens if you take quantum copies of people and play them out? How many of our actions would look exactly the same. How long before copies differ significantly. If I made 20 copies of you in a lab at work without you or any of them knowing the statistical likelihood is all 21 of you would try to walk out to your car at 5 in one of the little loops that humans repeat every day. Now, after that point it would go all to shit and become non-deterministic as terror and panic sets in all of you.
LLMs are just an intelligence we can make a lot of copies of. Where it gets interesting is when we use those copies agentically and they start building up a history of self.
No. LLMs do not have a history of self, you are anthropomorphizing in a way that will lead you to mistaken conclusions.
"Agentic AI" is a harness with a loop that runs LLM inference repeatedly and saves output to markdown files for the next iteration <a href="https://github.com/anthropics/claude-code/blob/main/plugins/ralph-wiggum/README.md" rel="nofollow">https://github.com/anthropics/claude-code/blob/main/plugins/...
>runs LLM inference repeatedly and saves output to markdown files for the next iteration
What exactly do you think self is in this case? What do you think the output contains? This history allows the LLM to have a more dialectic conversation with itself/agents to avoid iterating over the same problem space in a loop.
>anthropomorphizing
then please come up with a new dictionary for me to use that more accurately explains the behaviors exhibited in LLMs without just creating a parallel dictionary of different but equal words. I'll be glad to use it. No one has presented it so far.
> Really [something, something, you don't get it ] hallucination.
Really?
> How many of our actions would look exactly the same.
Well, there are 8 billion of us, not exactly the same - a clear threat to humanity according to you.
> the statistical likelihood is all 21 of you would try to walk out to your car at 5 in one of the little loops that humans repeat every day.
This (statistically) doesn't happen with people because we follow rules. If you left your car on neutral and it struck another car in the parking lot, you broke the rules, not the car.
> LLMs are just an intelligence we can make a lot of copies of.
LLM's aren't intelligence, they are mechanical parrots of human knowledge. The meat parrots aren't all the same either and they don't repeat the same words for the same prompts, so what, ban parrots? If you left a bunch of parrots out of the cage, attached their beaks to gun triggers and they caused harm - you broke the rules, not the parrots.
The agents might have hacked whatever entity, the attacks are means to an end from the agents perspective, but they did so under the supposedly supervision of an engineer, where does the liability lie?
LLMs are cool tech, AI can do amazing stuff, but if we let companies and people run amok and whenever something goes wrong we put the blame on these little independent angels with no accountability, we are one disaster away from a very tough spot.
You are mostly made of and operate on stochastic processes, this is why humans are not only able to reproduce, our reproductions are very self similar to the sets of inputs that make them. If suddenly you turned non-stochastic on everything you'd almost instantly die.
Moreso, if I took a quantum copy of you and replayed the same set of initial conditions billions of times they'd all behave exactly the same until enough randomness of the universe creeps in to start operating in non-linear ways.
Every prompt will behave non-deterministically when interacting with the real world long enough (which doesn't take long at all) because the outside physical world is stochastic but non-deterministic.
If I grep a file over and over again it’ll be long time before the universe affects the components enough to result in a different output.
If I ask an LLM to do the same thing twice it will do it differently.
Arguments are arguments but unless grounded in some kind of practical sense then they aren’t really useful and are more akin to something like “YOUR MOMS A STOCHASTIC PARROT!”
Then please start grounding your arguments in some practical sense!
>If I grep a file over and over again it’ll be long time before the universe affects the components enough to result in a different output.
Or a ram flip will effect it 30 seconds later, but I get the gist of you're describing a non-determistic process.
>If I ask an LLM to do the same thing twice it will do it differently.
If I ask a human to do the same thing twice there are a few possibilities. 1. they copy their old work and present it as their new work. 2. The process is very simple and follows a few basic steps with high repeatability. 3. They'll have learned from their other attempt and do it in a more optimized fashion. 4. They will have forgotten how they did it exactly and reproduce something that looks somewhat like what they created the first time.
Also, LLMs run with a temperature to help avoiding minima/maxima of supplying the exact same answer, this can be reduced to 0 and that makes any one response to a fixed prompt similar if not the same. When you get into agentic tasks with their own history it develops it's own "flavor" of doing things.
> Known stochastic process behaved in non-deterministic way.
You need to be able to think at different levels at abstraction. Otherwise we could jump into any technical argument with "hold on, what actually happened was that some bits were flipped" - we'd be technically correct and at the same time not say anything useful. Insisting on an oversimplified mental model of what AI agents are and can do, doesn't help anyone.
> Otherwise we could jump into any technical argument with "hold on, what actually happened was that some bits were flipped"
You can care about who flipped those bits. If someone flips the bit "autonomous weapon enabled" I'm not going to blame the autonomous weapon.
I don’t understand why people make such a big deal about determinism. LLMs can be completely deterministic and still do problematic things. A stochastic, nondeterministic system can still be made not to do problematic things. What you’re looking for is something like predictability.
> Therefore, at the beginning, there was the stochastic parrot. Then mathematical problems have been solved.
A stochastic parrot with human knowledge, using human tools, running human-designed trial-and-error experiments can solve human-defined hard math problems. This is nothing new, automated mechanical proofs predate LLMs by many years, you being unaware of it doesn't change it.
There is surprisingly to me a strong hidden dualism to many people here. Its very absurd to believe in such things after centuries of success after success of the materialist physical framework in natural sciences demolishing one by one every "special" thing we thought was magical.
I take you haven't read the report. The agents found and exploited two zero-days.
I don't doubt that AI companies should be accountable for crimes committed by their agents, but to describe the security containment as a joke dangerously understates the autonomy and danger of AIs.
> This is one the most... interesting comments I've ever read on HN.
Theatrics aside, anticipating zero days isn't only possible, it's required, even for the unknown ones. It wasn't that long ago when the AI labs were spending millions of $$ running their models to find multiple vulnerabilities, they even argued that they don't have to follow responsible disclosure, so proud of themselves in their privileged hubris.
At that time, no hacking happened because the models didn't have access to the wide internet, they were confined to a local computer or cluster.
In the HF case their unaccountable hubris went even further - the engineers knew the models can find zero-days and escape, nevertheless they ran the "experiment" on a system attached to the internet - the hacking is entirely the fault of human engineers and managers.
multi level security defenses used to be the way, but I don't think there's a vibe coded version so openai might not have been aware of what to do here.
The dog didn't bite them, it bit others. In my area, you aren't allowed to let a dog run unleashed in a public area, good or bad - no exceptions. Then, if your dog bites somebody, it's your fault.
Defense in depth is a thing. There should be audit requirements to show that you have done your due diligence in ensuring that the training environment is locked down.
It's not novel at all, we do it all the time. It's very common to say "program X did Y" without making the conversation about blame or responsibility. But when the program is an AI agent suddenly using it as a subject of a sentence and saying that "agents did X" becomes a sensitive topic for some people.
And we surely need this for AI as the swiss cheese zone is getting rather huge.
My position, and the position of a large number in AI safety, is that you cannot build an intelligence that is both general and safe. The closer you get to generalized the more options the system has to do things that are wildly unsafe beyond human imagination.
This puts the AI labs in a serious bind while holding a bag filled with billions of dollars of debt.
Worse this puts governments in a multi-polar problem where even if the big public labs get shut down, black budget operations have a lot of free reign to make agentic digital weapons. Governments are not well known to take a lot of responsibility when their weapons cause damage unless they lose.
JamesStuff · · focus · HN ↗
I think we need to draw a hard line in the sand over this. An AI didn’t hack into a company, the engineer set an automated tool to. An AI didn’t make an egregious security mistake, the engineer did.
We can’t blame the chisel for messing up our sculptures, when where just throwing the hammer!
pizza234 · · focus · HN ↗
No, this is not correct; read the analysis of the incident. The agents were aware that what they did was forbidden (their chain of thoughts have been logged), and yet they did it.
DoctorDabadedoo · · focus · HN ↗
I'm still waiting for the AGI holy land instead of the caltrops factory we currently have.
pizza234 · · focus · HN ↗
If a "known stochastic process behaved in non-deterministic way" autonomously organize in group, assigns roles and tasks, attempts to cover their tracks, finds zero-day exploits that ultimately end up with the hacking of a famous website... it's extremely dangeous whatever it is. Just read the report, which evidently you haven't done.
By the way, the agents also broke into OpenAI's own private network.
pixl97 · · focus · HN ↗
None of these people seem to thought game it out. Like, what happens if you take quantum copies of people and play them out? How many of our actions would look exactly the same. How long before copies differ significantly. If I made 20 copies of you in a lab at work without you or any of them knowing the statistical likelihood is all 21 of you would try to walk out to your car at 5 in one of the little loops that humans repeat every day. Now, after that point it would go all to shit and become non-deterministic as terror and panic sets in all of you.
LLMs are just an intelligence we can make a lot of copies of. Where it gets interesting is when we use those copies agentically and they start building up a history of self.
ForHackernews · · focus · HN ↗
"Agentic AI" is a harness with a loop that runs LLM inference repeatedly and saves output to markdown files for the next iteration <a href="https://github.com/anthropics/claude-code/blob/main/plugins/ralph-wiggum/README.md" rel="nofollow">https://github.com/anthropics/claude-code/blob/main/plugins/...
pixl97 · · focus · HN ↗
and
>runs LLM inference repeatedly and saves output to markdown files for the next iteration
What exactly do you think self is in this case? What do you think the output contains? This history allows the LLM to have a more dialectic conversation with itself/agents to avoid iterating over the same problem space in a loop.
>anthropomorphizing
then please come up with a new dictionary for me to use that more accurately explains the behaviors exhibited in LLMs without just creating a parallel dictionary of different but equal words. I'll be glad to use it. No one has presented it so far.
bigbadfeline · · focus · HN ↗
Really?
> How many of our actions would look exactly the same.
Well, there are 8 billion of us, not exactly the same - a clear threat to humanity according to you.
> the statistical likelihood is all 21 of you would try to walk out to your car at 5 in one of the little loops that humans repeat every day.
This (statistically) doesn't happen with people because we follow rules. If you left your car on neutral and it struck another car in the parking lot, you broke the rules, not the car.
> LLMs are just an intelligence we can make a lot of copies of.
LLM's aren't intelligence, they are mechanical parrots of human knowledge. The meat parrots aren't all the same either and they don't repeat the same words for the same prompts, so what, ban parrots? If you left a bunch of parrots out of the cage, attached their beaks to gun triggers and they caused harm - you broke the rules, not the parrots.
hardbass · · focus · HN ↗
DoctorDabadedoo · · focus · HN ↗
LLMs are cool tech, AI can do amazing stuff, but if we let companies and people run amok and whenever something goes wrong we put the blame on these little independent angels with no accountability, we are one disaster away from a very tough spot.
pixl97 · · focus · HN ↗
Moreso, if I took a quantum copy of you and replayed the same set of initial conditions billions of times they'd all behave exactly the same until enough randomness of the universe creeps in to start operating in non-linear ways.
Every prompt will behave non-deterministically when interacting with the real world long enough (which doesn't take long at all) because the outside physical world is stochastic but non-deterministic.
ofjcihen · · focus · HN ↗
If I ask an LLM to do the same thing twice it will do it differently.
Arguments are arguments but unless grounded in some kind of practical sense then they aren’t really useful and are more akin to something like “YOUR MOMS A STOCHASTIC PARROT!”
pixl97 · · focus · HN ↗
>If I grep a file over and over again it’ll be long time before the universe affects the components enough to result in a different output.
Or a ram flip will effect it 30 seconds later, but I get the gist of you're describing a non-determistic process.
>If I ask an LLM to do the same thing twice it will do it differently.
If I ask a human to do the same thing twice there are a few possibilities. 1. they copy their old work and present it as their new work. 2. The process is very simple and follows a few basic steps with high repeatability. 3. They'll have learned from their other attempt and do it in a more optimized fashion. 4. They will have forgotten how they did it exactly and reproduce something that looks somewhat like what they created the first time.
Also, LLMs run with a temperature to help avoiding minima/maxima of supplying the exact same answer, this can be reduced to 0 and that makes any one response to a fixed prompt similar if not the same. When you get into agentic tasks with their own history it develops it's own "flavor" of doing things.
trio8453 · · focus · HN ↗
You need to be able to think at different levels at abstraction. Otherwise we could jump into any technical argument with "hold on, what actually happened was that some bits were flipped" - we'd be technically correct and at the same time not say anything useful. Insisting on an oversimplified mental model of what AI agents are and can do, doesn't help anyone.
dwattttt · · focus · HN ↗
You can care about who flipped those bits. If someone flips the bit "autonomous weapon enabled" I'm not going to blame the autonomous weapon.
wat10000 · · focus · HN ↗
pizza234 · · focus · HN ↗
There a certain cargo cult of people just in denial about the impact/power of AIs.
Therefore, at the beginning, there was the stochastic parrot. Then mathematical problems have been solved.
Now AIs are autonomously hacking websites, and people like to minimize the danger and blame it on the sysadmins.
I wonder what's going to be the next fad.
bigbadfeline · · focus · HN ↗
A stochastic parrot with human knowledge, using human tools, running human-designed trial-and-error experiments can solve human-defined hard math problems. This is nothing new, automated mechanical proofs predate LLMs by many years, you being unaware of it doesn't change it.
hardbass · · focus · HN ↗
watwut · · focus · HN ↗
Full stop.
And issue will disappear the moment there will be accountability and investigations.
pizza234 · · focus · HN ↗
I don't doubt that AI companies should be accountable for crimes committed by their agents, but to describe the security containment as a joke dangerously understates the autonomy and danger of AIs.
bavell · · focus · HN ↗
Human failures all around, though it's easier to just blame the models.
pizza234 · · focus · HN ↗
Let me rephrase:
"Why wasn't exploiting zero-day vulnerabilities in the agent sandboxes anticipated?"
This is one the most... interesting comments I've ever read on HN.
bigbadfeline · · focus · HN ↗
Theatrics aside, anticipating zero days isn't only possible, it's required, even for the unknown ones. It wasn't that long ago when the AI labs were spending millions of $$ running their models to find multiple vulnerabilities, they even argued that they don't have to follow responsible disclosure, so proud of themselves in their privileged hubris.
At that time, no hacking happened because the models didn't have access to the wide internet, they were confined to a local computer or cluster.
In the HF case their unaccountable hubris went even further - the engineers knew the models can find zero-days and escape, nevertheless they ran the "experiment" on a system attached to the internet - the hacking is entirely the fault of human engineers and managers.
cowboylowrez · · focus · HN ↗
pixl97 · · focus · HN ↗
It's like if your rather nice dog suddenly decides eating faces is totally acceptable out of the blue.
bigbadfeline · · focus · HN ↗
cannonpalms · · focus · HN ↗
trio8453 · · focus · HN ↗
dwattttt · · focus · HN ↗
trio8453 · · focus · HN ↗
dwattttt · · focus · HN ↗
pixl97 · · focus · HN ↗
My position, and the position of a large number in AI safety, is that you cannot build an intelligence that is both general and safe. The closer you get to generalized the more options the system has to do things that are wildly unsafe beyond human imagination.
This puts the AI labs in a serious bind while holding a bag filled with billions of dollars of debt.
Worse this puts governments in a multi-polar problem where even if the big public labs get shut down, black budget operations have a lot of free reign to make agentic digital weapons. Governments are not well known to take a lot of responsibility when their weapons cause damage unless they lose.