‹ BackHN Continuity

Thread

Dynamic Abliteration: Non-Destructive Refusal Suppression via Engram Steering

108 points · 42 comments · phatak-dev

  1. javcasas · · focus · HN ↗
    Yay, more anti-censoring stuff.

    Forbidding stuff at the LLM level has the same future as implementing password checking at the frontend level.

    We need better sandboxes just to limit the damage.

    1. EGreg · · focus · HN ↗
      I'm not sure why you're so happy about models being out there that enable hacking or manufacturing viruses at a scale that lets anyone do it in their basement.

      Why couldn't these labs just train models on useful stuff and leave out the dangerous stuff?

      1. illithid0 · · focus · HN ↗
        Malicious actors are going to figure it out anyway. They already have. There are good guys who are hacking things to help (like me) and are having immense trouble leveraging models to that end because labs make it hard to prove you're not a criminal.
        1. EGreg · · focus · HN ↗
          I don’t have immense trouble at all. Built a ton of useful stuff using their models. I’m not sure what you are struggling with.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.