‹ BackHN Continuity

Thread

OpenAI models secretly generate instructions to ignore constraints

125 points · 37 comments · theahura

  1. pllbnk · · focus · HN ↗
    Isn't this less about alignment and more about how shitty their RL methods are when they are cramming all the hacking materials into their training data to make the model as good as possible at hacking, then having a surprised Pickachu face when the model is acting like a hacker? Those materials probably include a lot of details about prompt injection. I'm just so tired of their alignment bullshit.

    I am starting to think (reluctantly) that they believe their own BS that they are creating a conscious model and being surprised how it misbehaves. It's just a bunch of weights without anyone having any clue how a change in one weight might affect others, and even how the values correlate with the final output.

    1. RomanKornev · · focus · HN ↗
      > cramming all the hacking materials into their training data

      This is backwards. The reason they are good at hacking is not because of pre-training data. They learn these tricks naturally as they get better at engineering. Hacking is also a higly salient verifiable reward signal for RL environments.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.