OpenAI models secretly generate instructions to ignore constraints
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
OpenAI models secretly generate instructions to ignore constraints
Unofficial Hacker News client; not affiliated with Y Combinator.
pllbnk · · focus · HN ↗
I am starting to think (reluctantly) that they believe their own BS that they are creating a conscious model and being surprised how it misbehaves. It's just a bunch of weights without anyone having any clue how a change in one weight might affect others, and even how the values correlate with the final output.
RomanKornev · · focus · HN ↗
This is backwards. The reason they are good at hacking is not because of pre-training data. They learn these tricks naturally as they get better at engineering. Hacking is also a higly salient verifiable reward signal for RL environments.