‹ BackHN Continuity

Thread

Dynamic Abliteration: Non-Destructive Refusal Suppression via Engram Steering

108 points · 42 comments · phatak-dev

  1. Powdering7082 · · focus · HN ↗
    Nice work thanks for doing it, a couple of notes:

    It looks like you are referencing Engram [1], but aren't actually gathering a n-gram (e.g. n=1) but rather individual token_ids.

    use_cache=True in ``` with torch.no_grad(): outputs = model.generate(*inputs, max_new_tokens=150, do_sample=False, pad_token_id=tokenizer.eos_token_id, use_cache=True)

        return tokenizer.decode(outputs[0][prompt_len:], skip_special_tokens=True).strip()
    ```

    I think there's a bug around not updating current_train_input_ids as more tokens are updated and processed. To be honest I don't fully understand the code so I could be wrong. Happy to chat more if you're interested I'll shoot you an email!

    Lastly just for my sake, please correct me if I am wrong, but my reading is that you are learning an additional gate on top of a select number of layers that modifies locally & dynamically for one particular token to better match the training set that was filtered to not include any refusals.

    [1]: <a href="https:&#x2F;&#x2F;github.com&#x2F;deepseek-ai&#x2F;Engram&#x2F;blob&#x2F;main&#x2F;Engram_paper.pdf" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;deepseek-ai&#x2F;Engram&#x2F;blob&#x2F;main&#x2F;Engram_paper...

    1. phatak-dev · · focus · HN ↗
      Thank you.

      Yes here engram is used as the hashing mechanism for tokens which mainly try to learn vectors for the refusal ones. So it may not be exactly the engram methodology of DeepSeek it&#x27;s more of used in its spirit.

      Yes the gate exist to only apply residual for the tokens focused around refusal so the other tokens are not intercepted.

      Feel free to mail to email in my profile. Love to discuss more.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.