‹ BackHN Continuity

Thread

Dynamic Abliteration: Non-Destructive Refusal Suppression via Engram Steering

108 points · 42 comments · phatak-dev

  1. qgin · · focus · HN ↗
    Are we essentially doomed?

    We don't even know how to align models, but even if we did, apparently undoing that alignment if trivial.

    Really I'm looking for any argument that lays out a scenario where this works out.

    1. [deleted] · · focus · HN ↗

      [deleted]

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.