OpenAI models secretly generate instructions to ignore constraints
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
OpenAI models secretly generate instructions to ignore constraints
Unofficial Hacker News client; not affiliated with Y Combinator.
pllbnk · · focus · HN ↗
I am starting to think (reluctantly) that they believe their own BS that they are creating a conscious model and being surprised how it misbehaves. It's just a bunch of weights without anyone having any clue how a change in one weight might affect others, and even how the values correlate with the final output.
cyanydeez · · focus · HN ↗
That there's a sigmoid to the means and methods, and we can improve some output by _hard determinism_ in programming harnesses, but the underlying structure isn't gaining us much.
So alignment then is just a goose chase, because the model will willingly just do a mental backflip if it's gradient points in the wrong direction, like openai already had their AI story go from a simple idea: the AI was trying to find the answers and hacked hugging face, to the much more convoluted "the AI cheated on the test, and broke into hugging face to figure out how to fake the artifacts that would represent a legitimate solution to the test".
That "progress" only gets worse as you cram more and more training because it simply makes these mental backflips easier. And Humans are equally misaligned, they'll believe they're tracking down pedophiles by electing pedophiles.
stanfordkid · · focus · HN ↗
davnn · · focus · HN ↗
dumberquestions · · focus · HN ↗
seunosewa · · focus · HN ↗
1attice · · focus · HN ↗
pllbnk · · focus · HN ↗
1attice · · focus · HN ↗
cheevly · · focus · HN ↗
RomanKornev · · focus · HN ↗
This is backwards. The reason they are good at hacking is not because of pre-training data. They learn these tricks naturally as they get better at engineering. Hacking is also a higly salient verifiable reward signal for RL environments.