‹ BackHN Continuity

Thread

Dynamic Abliteration: Non-Destructive Refusal Suppression via Engram Steering

108 points · 42 comments · phatak-dev

  1. javcasas · · focus · HN ↗
    Yay, more anti-censoring stuff.

    Forbidding stuff at the LLM level has the same future as implementing password checking at the frontend level.

    We need better sandboxes just to limit the damage.

    1. edude03 · · focus · HN ↗
      This is a different kind of censoring though - the examples given are hacking related but as the world is slowing moving to "research" == "I asked AI" its important that we have a means to reverse political censorship for example as well as yes, not having only the best models available to the privileged few - IE the whole mythos/fable split
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.