‹ BackHN Continuity

Thread

OpenAI models secretly generate instructions to ignore constraints

125 points · 37 comments · theahura

  1. pllbnk · · focus · HN ↗
    Isn't this less about alignment and more about how shitty their RL methods are when they are cramming all the hacking materials into their training data to make the model as good as possible at hacking, then having a surprised Pickachu face when the model is acting like a hacker? Those materials probably include a lot of details about prompt injection. I'm just so tired of their alignment bullshit.

    I am starting to think (reluctantly) that they believe their own BS that they are creating a conscious model and being surprised how it misbehaves. It's just a bunch of weights without anyone having any clue how a change in one weight might affect others, and even how the values correlate with the final output.

    1. 1attice · · focus · HN ↗
      you are just a bunch of neural weights in wetware. Pot, kettle. No lateral nn-on-nn violence pls
      1. pllbnk · · focus · HN ↗
        Oh please, I am a general intelligence, not a wannabe-AGI neural network. I can operate with a few orders of magnitude less power at much higher TPS and still call out a large portion of BS LLMs spit out.
        1. 1attice · · focus · HN ↗
          Sure, but can you help me with Navier Stokes
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.