‹ BackHN Continuity

Thread

OpenAI models secretly generate instructions to ignore constraints

125 points · 37 comments · theahura

  1. greatgib · · focus · HN ↗
    They obviously carefully avoid to give the whole context and original prompt instructions they gave to the model. Or model simulating the user part.

    I would easily guess that if you got it you will understand that this behavior is not natural but induced by the researcher.

    And to be noted in addition that they are standard prompt injections that were rejected anyway as such.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.