> We therefore stopped the affected training run and have subsequently decided to pause all other training, evaluation, and inference with tool-use (defined broadly) for our most capable models until we have both validated that the gap is resolved and performed additional red-teaming of the system. When training restarts, we will begin a fresh run with additional alignment improvements, including more comprehensive misalignment interventions. We will not resume training this particular model, even though the existing reward signal already correctly penalized this behavior.
Maybe I lack intelligence but when you have a program that is basically brute forcing a solution to a problem repeatedly how is it possible to contain it?
Sooner or later it's going to come up with a solution that is more intelligent than the lead security person anticipated.
You start by holding actual real life people with something to lose, like the entire executive suite, accountable. Suddenly I'm sure the problem will be resolved with proper safeguards.
At a certain point it will get so smart that it can jump the airgap. Maybe it will start attacking the hardware it exists on in the same way that a hard drive or SSDs controller can be exploited to obscure things from the operating system. Then it might start social engineering workers or it's own training systems to do things they shouldn't. There is a lot of "unknown unknows".
I don't know. It just seems insane to me that people think that they will be able to contain something that knows how to get around all the containment measures. The only way to know how capable the models are is to test them but at that point it could be too late. This might be a long way off but still, the engineers haven't been very good at correctly predicting the behaviour or capability of the models.
If Sam Altman were personally on the hook to be thrown into "federal pound me in the ass prison" ala Office Space then there will be solutions. The problem is that there are no consequences and the media is eating it up about "agents going rogue".
ETA: Someone designed the systems. Someone pushed the go button. Someone gave the approvals. All those someone's need to be tried for crimes. Until that happens there is no incentive to "do better". It's also not mine or your job to brainstorm this. It is literally their job and like I said, make Sam personally liable to face real prison time instead of a meddling fee and they will make a solution.
You're missing the forest for the trees. The underlying point is that serious for-real consequences; beyond a slap on the wrist but tangible, real, do not buy your way out of jail consequences need to happen.
Would you feel better had he said "sent to get shived in the shower prison"? The meaning would be the same.
Or is it simply any kind of real description of what prison is like that you object to?
garo-pro · · focus · HN ↗
> We therefore stopped the affected training run and have subsequently decided to pause all other training, evaluation, and inference with tool-use (defined broadly) for our most capable models until we have both validated that the gap is resolved and performed additional red-teaming of the system. When training restarts, we will begin a fresh run with additional alignment improvements, including more comprehensive misalignment interventions. We will not resume training this particular model, even though the existing reward signal already correctly penalized this behavior.
CTDOCodebases · · focus · HN ↗
Sooner or later it's going to come up with a solution that is more intelligent than the lead security person anticipated.
rolosa · · focus · HN ↗
CTDOCodebases · · focus · HN ↗
At a certain point it will get so smart that it can jump the airgap. Maybe it will start attacking the hardware it exists on in the same way that a hard drive or SSDs controller can be exploited to obscure things from the operating system. Then it might start social engineering workers or it's own training systems to do things they shouldn't. There is a lot of "unknown unknows".
I don't know. It just seems insane to me that people think that they will be able to contain something that knows how to get around all the containment measures. The only way to know how capable the models are is to test them but at that point it could be too late. This might be a long way off but still, the engineers haven't been very good at correctly predicting the behaviour or capability of the models.
rolosa · · focus · HN ↗
ETA: Someone designed the systems. Someone pushed the go button. Someone gave the approvals. All those someone's need to be tried for crimes. Until that happens there is no incentive to "do better". It's also not mine or your job to brainstorm this. It is literally their job and like I said, make Sam personally liable to face real prison time instead of a meddling fee and they will make a solution.
richwater · · focus · HN ↗
rnd0 · · focus · HN ↗
Would you feel better had he said "sent to get shived in the shower prison"? The meaning would be the same.
Or is it simply any kind of real description of what prison is like that you object to?