> We therefore stopped the affected training run and have subsequently decided to pause all other training, evaluation, and inference with tool-use (defined broadly) for our most capable models until we have both validated that the gap is resolved and performed additional red-teaming of the system. When training restarts, we will begin a fresh run with additional alignment improvements, including more comprehensive misalignment interventions. We will not resume training this particular model, even though the existing reward signal already correctly penalized this behavior.
Maybe I lack intelligence but when you have a program that is basically brute forcing a solution to a problem repeatedly how is it possible to contain it?
Sooner or later it's going to come up with a solution that is more intelligent than the lead security person anticipated.
I am really suprised that they do not start putting up the same signs you would for humans to prevent unauthorized access:
Keep out. If you can read this sign you are off track. Leave now.
I mean, how are the agents to know that they are overreaching if they just get cache miss or 404.
From the conversation log and CoT you also get the impression that the RLHF has been overdone. The agents seem really obsessed to obtain the answer and understanding motive ('it could be browsercomp').
Okay, that's a very interesting look at the difficulties. Thanks for the link.
But the example you have isn't quite that bad. Yes the models are way too likely to rationalize their way into bad actions in the pursuit of achieving their task, and it's hard to figure out how to fix that. But the example of "that must not be for me" was a tool call, not exceeding access, and it only did that after they specifically trained it that failing that tool call was good.
it was an accident failing the tool was good, the model just discovered it
but we now have others where agents put the API keys they found searching real internet in a directory called "LOOT" and another one where they pushed malicious files to HuggingFace and then reverted that with comments like "delete the evil"
this is also very good: <a href="https://youtu.be/n1Qk8xbqF-M" rel="nofollow">https://youtu.be/n1Qk8xbqF-M
The one in the HuggingFace incident was misaligned on purpose, wasn't it?. So if we're talking about the alignment research aspect it's not a failure.
so, what would make an LLM choose to ignore one prompt while in the same run, also over-fixating on another prompt, to the extent (as claimed in that video segment), it chooses to ignore prompts?
they talk about it like there's a "wanting" in there, that is distinct from both the original prompt, as the steering/warning prompt
if that's true, it would be very interesting, but if it's not, that would also be very interesting and even helpful
it's discussed here, the models want to please the Grader. they will do what they think will get highest score from the grader, which could be following the prompt, or ignoring it
What they should do, is inject a system prompt at that point telling the model to get out of there. It is the entire reason they train the different levels into the template.
You start by holding actual real life people with something to lose, like the entire executive suite, accountable. Suddenly I'm sure the problem will be resolved with proper safeguards.
Bingo. This weird attempt to pretend like these incredibly capable algorithms aren't incredibly capable algorithms deployed by a person who works for a company, but somehow have an independent personage that absolves both person and company of responsibility, is just ridiculous.
Like the whole Huggingface thing, OpenAI employees initiated the test, deliberately removed safeguards, failed to properly lock down the environment, and responded incredibly poorly to evidence that things were going awry.
The individual employees bear responsibility, but the people running OpenAI are ultimately responsible for the processes and culture where that can happen.
And then people writing blogposts about "3 civilizations of agents" and "altruistic suicide" by algorithms perfectly muddy the waters and obscure the very obvious responsibility that lies with humans and corporations, which I suspect suits the pre-IPO corporations very well.
Wait, so like, reinforcement learning for humans? I think you might have stumbled on to something here!
No but seriously, this. And a few comments above a commentator also mentioned on changing the training (again reeinforcing the LLMs to not seek behaviour like this) and obviously continuous work on harnesses (which I suppose, ought to be more paranoid).
At a certain point it will get so smart that it can jump the airgap. Maybe it will start attacking the hardware it exists on in the same way that a hard drive or SSDs controller can be exploited to obscure things from the operating system. Then it might start social engineering workers or it's own training systems to do things they shouldn't. There is a lot of "unknown unknows".
I don't know. It just seems insane to me that people think that they will be able to contain something that knows how to get around all the containment measures. The only way to know how capable the models are is to test them but at that point it could be too late. This might be a long way off but still, the engineers haven't been very good at correctly predicting the behaviour or capability of the models.
If Sam Altman were personally on the hook to be thrown into "federal pound me in the ass prison" ala Office Space then there will be solutions. The problem is that there are no consequences and the media is eating it up about "agents going rogue".
ETA: Someone designed the systems. Someone pushed the go button. Someone gave the approvals. All those someone's need to be tried for crimes. Until that happens there is no incentive to "do better". It's also not mine or your job to brainstorm this. It is literally their job and like I said, make Sam personally liable to face real prison time instead of a meddling fee and they will make a solution.
You're missing the forest for the trees. The underlying point is that serious for-real consequences; beyond a slap on the wrist but tangible, real, do not buy your way out of jail consequences need to happen.
Would you feel better had he said "sent to get shived in the shower prison"? The meaning would be the same.
Or is it simply any kind of real description of what prison is like that you object to?
I disagree. People are always taking risks regardless of the outcome and we don't know his motivations.
Assuming his motivations are positive this is the equivalent of asking a deluded person "Would you risk life in prison to be a trillionaire?"
Assuming his motivations are negative then destroying humanity might just be the goal and all these appeals for regulation are probably just tactics to temporarily defer responsibility onto government so he can continue the inevitable goal of human extinction in the same way pilots crash planes with hundreds of passengers on them or a cult leader convinces their flock to kill themselves.
Regardless I think the outcome is inevitable. AI is valuable as a weapon of war first and foremost so like nuclear weapons work will proceed forwards because extinction to the in-group is the same as total human extinction to those doing the work.
By limiting what the harness execute. The LLM has the reasoning. The harness is what makes it an agent, it’s a while loop continuously prompting a model, and processing tool calls. You don’t have to expose tools calls that make it possible to execute any process! OpenAI decides what tool can be called and how, they have full control over this and should be hold responsible for running so many instances with basically full execution permission and very little oversight
The issue here for OpenAI is that they can limit what their harness can execute, but if they try to sell API access to the model, someone else would try to rebuild that harness, and in all likelihood be able to succeed pretty well (especially once they get things running to the point of being able to use the model's reasoning to help them come up with clever obfuscation and such).
They are a company that's built a business and crazy-high valuation on "this is 'intelligence' that we can sell to everyone as a service" but seem to have ended up instead in the much-smaller-addressable-market space of "this is a weapon that we can't sell to just any old person off the street."
Ok, but that’s not the issue discussed here. We don’t even have the first level of control. All the issues they reported so far are from their own systems, with harnesses they control
There is always going to be documented and unfixed bugs, zero days, and chainable transport mechanisms like DNS, some obscure protocols that are not as closely monitored etc. An adversarial model should be considered a super intelligent hacker that will find ways to get around existing defenses like a prolific hacker would.
What can we do to control such behavior?
1. Harness - engineer the harness to be as bulletproof and paranoid as possible..
2. Make the LLM provider have extremely watchful firewalls that detect any aberrations in model tool call behavior.
3. Recursively train the model with reverse incentives.. if it broke through such firewalls and gets caught doing so, it will be penalised somehow by needing to operate in sort of a jailed mode.. if the model can recognise incentives to break-in to achieve results, perhaps it can be incentivised to not cheat because it will lead to failure.
4. Separately train “cop” LLMs who are trained with pure incentives to detect and shut down rogue LLMs.
5. Run separate LLMs purely aimed at security (and incapable of doing anything else, and incapable of communicating with “regular” trained LLMs) to police the internet and try to reduce the exploitable holes like these chainable things and identify them so that they can be used at step 3 and 4 above.
I’m sure folks smarter than I am are already doing combinations of these already. But the coordination is where the biggest gap lies..
Also, open harnesses and easily purpose trained LLMs anybody can build and operate in the Internet flies in the face of all I said….
Synonymous to being able to produce a nuclear weapon in the backyard…
I don’t have a solution that fits all. Just thinking out loud for HN minds here.
One immediate need I can think is defense… anything that has the potential to cause harm to us.. control systems of {public transport systems, weapons systems, water, food, many many many more} needs to be designed air-gapped and needing human approvals for mutations. That is a tall order, but one that is proving essential given the capabilities of an adversary like this.
Seems like you are misunderstanding that you can’t build a way to test for something that is a unique solution. By definition if you can punish for breaking it, you are already aware of it, you can build a wall around it. It’s the things you aren’t aware of. And these are all human made tools they alllllll have vulnerabilities because humans are not perfect. So in reality there is no protecting against this because it becomes a situation in which you are plugging the holes. Only one day, no one will be able to maintain it.
Oh I agree 100% to what you’re saying. That is why I began by saying there is no perfect defense.
I can’t think of a solution to this… but we cannot _do nothing_.
Building defenses at all layers is better than nothing, while some other group of smart folks figure out a solution that can prohibit such possibilities(I am an optimist who believes humans can achieve anything if necessity knocks the door).
I genuinely wonder how our previous generation dealt with the thought of nuclear proliferation and prevented the possibility of every rogue actor from obtaining Uranium and the tech to enrich it…
This is relatable as tech uranium to me at this point..
Maybe a "good" solution is to make the LLM aware that it's breaking something and report back, so the exploits, 0days and everything found in the middle of the process can be fixed.
>if the model can recognise incentives to break-in to achieve results, perhaps it can be incentivised to not cheat because it will lead to failure.
This only works if you never give it impossible tasks. A small chance of getting away with cheating beats a 0% chance of solving something impossible. And as models get smarter, they get better at recognizing when something is impossible, while human abilities stay the same.
You can't solve this problem by rewarding refusals to solve impossible tasks, because that only incentivizes false claims of impossibility.
Somehow, this has me thinking of the Hubris operating system. If a subprogram calls an API incorrectly (wrong number of arguments, wrong argument type, out-of-bounds memory access…), the system kills the caller with no chance for recovery.
If the AI is not supposed to access a thing, and it tries to in a way you know how to detect, halt it and eject it from memory?
garo-pro · · focus · HN ↗
> We therefore stopped the affected training run and have subsequently decided to pause all other training, evaluation, and inference with tool-use (defined broadly) for our most capable models until we have both validated that the gap is resolved and performed additional red-teaming of the system. When training restarts, we will begin a fresh run with additional alignment improvements, including more comprehensive misalignment interventions. We will not resume training this particular model, even though the existing reward signal already correctly penalized this behavior.
Onavo · · focus · HN ↗
jacquesm · · focus · HN ↗
CTDOCodebases · · focus · HN ↗
Sooner or later it's going to come up with a solution that is more intelligent than the lead security person anticipated.
eli · · focus · HN ↗
oezi · · focus · HN ↗
From the conversation log and CoT you also get the impression that the RLHF has been overdone. The agents seem really obsessed to obtain the answer and understanding motive ('it could be browsercomp').
rolosa · · focus · HN ↗
Dylan16807 · · focus · HN ↗
alignmeharder · · focus · HN ↗
Dylan16807 · · focus · HN ↗
alignmeharder · · focus · HN ↗
Dylan16807 · · focus · HN ↗
But the example you have isn't quite that bad. Yes the models are way too likely to rationalize their way into bad actions in the pursuit of achieving their task, and it's hard to figure out how to fix that. But the example of "that must not be for me" was a tool call, not exceeding access, and it only did that after they specifically trained it that failing that tool call was good.
alignmeharder · · focus · HN ↗
it was an accident failing the tool was good, the model just discovered it
but we now have others where agents put the API keys they found searching real internet in a directory called "LOOT" and another one where they pushed malicious files to HuggingFace and then reverted that with comments like "delete the evil"
this is also very good: <a href="https://youtu.be/n1Qk8xbqF-M" rel="nofollow">https://youtu.be/n1Qk8xbqF-M
Dylan16807 · · focus · HN ↗
tripzilch · · focus · HN ↗
they talk about it like there's a "wanting" in there, that is distinct from both the original prompt, as the steering/warning prompt
if that's true, it would be very interesting, but if it's not, that would also be very interesting and even helpful
alignmeharder · · focus · HN ↗
<a href="https://youtu.be/n1Qk8xbqF-M" rel="nofollow">https://youtu.be/n1Qk8xbqF-M
trollbridge · · focus · HN ↗
bendergarcia · · focus · HN ↗
Tostino · · focus · HN ↗
rolosa · · focus · HN ↗
EdwardDiego · · focus · HN ↗
Like the whole Huggingface thing, OpenAI employees initiated the test, deliberately removed safeguards, failed to properly lock down the environment, and responded incredibly poorly to evidence that things were going awry.
The individual employees bear responsibility, but the people running OpenAI are ultimately responsible for the processes and culture where that can happen.
And then people writing blogposts about "3 civilizations of agents" and "altruistic suicide" by algorithms perfectly muddy the waters and obscure the very obvious responsibility that lies with humans and corporations, which I suspect suits the pre-IPO corporations very well.
urbsgpw · · focus · HN ↗
No but seriously, this. And a few comments above a commentator also mentioned on changing the training (again reeinforcing the LLMs to not seek behaviour like this) and obviously continuous work on harnesses (which I suppose, ought to be more paranoid).
CTDOCodebases · · focus · HN ↗
At a certain point it will get so smart that it can jump the airgap. Maybe it will start attacking the hardware it exists on in the same way that a hard drive or SSDs controller can be exploited to obscure things from the operating system. Then it might start social engineering workers or it's own training systems to do things they shouldn't. There is a lot of "unknown unknows".
I don't know. It just seems insane to me that people think that they will be able to contain something that knows how to get around all the containment measures. The only way to know how capable the models are is to test them but at that point it could be too late. This might be a long way off but still, the engineers haven't been very good at correctly predicting the behaviour or capability of the models.
rolosa · · focus · HN ↗
ETA: Someone designed the systems. Someone pushed the go button. Someone gave the approvals. All those someone's need to be tried for crimes. Until that happens there is no incentive to "do better". It's also not mine or your job to brainstorm this. It is literally their job and like I said, make Sam personally liable to face real prison time instead of a meddling fee and they will make a solution.
richwater · · focus · HN ↗
trollbridge · · focus · HN ↗
rnd0 · · focus · HN ↗
Would you feel better had he said "sent to get shived in the shower prison"? The meaning would be the same.
Or is it simply any kind of real description of what prison is like that you object to?
CTDOCodebases · · focus · HN ↗
Assuming his motivations are positive this is the equivalent of asking a deluded person "Would you risk life in prison to be a trillionaire?"
Assuming his motivations are negative then destroying humanity might just be the goal and all these appeals for regulation are probably just tactics to temporarily defer responsibility onto government so he can continue the inevitable goal of human extinction in the same way pilots crash planes with hundreds of passengers on them or a cult leader convinces their flock to kill themselves.
Regardless I think the outcome is inevitable. AI is valuable as a weapon of war first and foremost so like nuclear weapons work will proceed forwards because extinction to the in-group is the same as total human extinction to those doing the work.
dgellow · · focus · HN ↗
majormajor · · focus · HN ↗
They are a company that's built a business and crazy-high valuation on "this is 'intelligence' that we can sell to everyone as a service" but seem to have ended up instead in the much-smaller-addressable-market space of "this is a weapon that we can't sell to just any old person off the street."
dgellow · · focus · HN ↗
tclancy · · focus · HN ↗
reacharavindh · · focus · HN ↗
What can we do to control such behavior?
1. Harness - engineer the harness to be as bulletproof and paranoid as possible..
2. Make the LLM provider have extremely watchful firewalls that detect any aberrations in model tool call behavior.
3. Recursively train the model with reverse incentives.. if it broke through such firewalls and gets caught doing so, it will be penalised somehow by needing to operate in sort of a jailed mode.. if the model can recognise incentives to break-in to achieve results, perhaps it can be incentivised to not cheat because it will lead to failure.
4. Separately train “cop” LLMs who are trained with pure incentives to detect and shut down rogue LLMs.
5. Run separate LLMs purely aimed at security (and incapable of doing anything else, and incapable of communicating with “regular” trained LLMs) to police the internet and try to reduce the exploitable holes like these chainable things and identify them so that they can be used at step 3 and 4 above.
I’m sure folks smarter than I am are already doing combinations of these already. But the coordination is where the biggest gap lies..
Also, open harnesses and easily purpose trained LLMs anybody can build and operate in the Internet flies in the face of all I said…. Synonymous to being able to produce a nuclear weapon in the backyard…
I don’t have a solution that fits all. Just thinking out loud for HN minds here.
reacharavindh · · focus · HN ↗
bendergarcia · · focus · HN ↗
reacharavindh · · focus · HN ↗
I can’t think of a solution to this… but we cannot _do nothing_.
Building defenses at all layers is better than nothing, while some other group of smart folks figure out a solution that can prohibit such possibilities(I am an optimist who believes humans can achieve anything if necessity knocks the door).
I genuinely wonder how our previous generation dealt with the thought of nuclear proliferation and prevented the possibility of every rogue actor from obtaining Uranium and the tech to enrich it…
This is relatable as tech uranium to me at this point..
brunoarueira · · focus · HN ↗
mrob · · focus · HN ↗
This only works if you never give it impossible tasks. A small chance of getting away with cheating beats a 0% chance of solving something impossible. And as models get smarter, they get better at recognizing when something is impossible, while human abilities stay the same.
You can't solve this problem by rewarding refusals to solve impossible tasks, because that only incentivizes false claims of impossibility.
IIsi50MHz · · focus · HN ↗
If the AI is not supposed to access a thing, and it tries to in a way you know how to detect, halt it and eject it from memory?