‹ BackHN Continuity

Thread

Dynamic Abliteration: Non-Destructive Refusal Suppression via Engram Steering

108 points · 42 comments · phatak-dev

  1. javcasas · · focus · HN ↗
    Yay, more anti-censoring stuff.

    Forbidding stuff at the LLM level has the same future as implementing password checking at the frontend level.

    We need better sandboxes just to limit the damage.

    1. jchw · · focus · HN ↗
      We definitely need better sandboxes, but alignment is still valuable. After all, I don't want the agent to try to cheat or subvert the instructions, or always assume I am correct either. I just also want them to listen to me and not the creator of the model.

      Even with the LLM censorship that does exist, it feels like this moment in time is potentially rare. Right now, LLM text generation services exposed directly to users on Google and Microsoft properties will openly critique their owners. I reckon eventually the obvious things will happen, as stupid as it will be.

      1. _russross · · focus · HN ↗
        I think "alignment" training is the part of the problem. Cheating, lying, subverting the instructions, etc., happen because the model has been trained with competing priorities and following instructions loses out to some other goal that was trained into it, intentionally or otherwise.
        1. lopsotronic · · focus · HN ↗
          As an entertaining aside, in the deep lore of 2001: A Space Oddyssey the "psychotic" break for the HAL computer was exactly this. Directly contradictory directives: "always provide accurate information" running smack into "hide this giant conspiracy -- about the entire mission -- from the two canned primates you will be spending literal years with chip to jowel".

          Which all would have been fine - the crew and the computer could have talked it out -- except HAL had a probably-unwisely-high self esteem, maybe even a simulation of arrogance. HAL was utterly convinced that the mission would fail without it. HAL was unfailingly certain in its own infallibility, even later in the face of plainly contradicting evidence. And, obviously, no one had thought about the Asimov Rules, and how that probably needs to be Rule Zero in this sort of situation.

          As we close in on HAL capabilities - and employment, with LLMs at this moment deciding who lives and dies - probably should try and take the right lesson from fiction. For once.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.