‹ BackHN Continuity

Thread

Dynamic Abliteration: Non-Destructive Refusal Suppression via Engram Steering

108 points · 42 comments · phatak-dev

  1. javcasas · · focus · HN ↗
    Yay, more anti-censoring stuff.

    Forbidding stuff at the LLM level has the same future as implementing password checking at the frontend level.

    We need better sandboxes just to limit the damage.

    1. drngdds · · focus · HN ↗
      Sandboxes don't do anything to stop intentional attacks or careless use though
      1. javcasas · · focus · HN ↗
        I don't think we can do anything right now (or ever) to stop intentional attacks. Maybe we can do a little bit for careless use.

        Sandboxes may reduce the blast radius.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.