‹ BackHN Continuity

Thread

Dynamic Abliteration: Non-Destructive Refusal Suppression via Engram Steering

108 points · 42 comments · phatak-dev

  1. javcasas · · focus · HN ↗
    Yay, more anti-censoring stuff.

    Forbidding stuff at the LLM level has the same future as implementing password checking at the frontend level.

    We need better sandboxes just to limit the damage.

    1. EGreg · · focus · HN ↗
      I'm not sure why you're so happy about models being out there that enable hacking or manufacturing viruses at a scale that lets anyone do it in their basement.

      Why couldn't these labs just train models on useful stuff and leave out the dangerous stuff?

      1. [deleted] · · focus · HN ↗

        [deleted]

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.