Dynamic Abliteration: Non-Destructive Refusal Suppression via Engram Steering
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Dynamic Abliteration: Non-Destructive Refusal Suppression via Engram Steering
Unofficial Hacker News client; not affiliated with Y Combinator.
Powdering7082 · · focus · HN ↗
It looks like you are referencing Engram [1], but aren't actually gathering a n-gram (e.g. n=1) but rather individual token_ids.
use_cache=True in ``` with torch.no_grad(): outputs = model.generate(*inputs, max_new_tokens=150, do_sample=False, pad_token_id=tokenizer.eos_token_id, use_cache=True)
```I think there's a bug around not updating current_train_input_ids as more tokens are updated and processed. To be honest I don't fully understand the code so I could be wrong. Happy to chat more if you're interested I'll shoot you an email!
Lastly just for my sake, please correct me if I am wrong, but my reading is that you are learning an additional gate on top of a select number of layers that modifies locally & dynamically for one particular token to better match the training set that was filtered to not include any refusals.
[1]: <a href="https://github.com/deepseek-ai/Engram/blob/main/Engram_paper.pdf" rel="nofollow">https://github.com/deepseek-ai/Engram/blob/main/Engram_paper...
phatak-dev · · focus · HN ↗
Yes here engram is used as the hashing mechanism for tokens which mainly try to learn vectors for the refusal ones. So it may not be exactly the engram methodology of DeepSeek it's more of used in its spirit.
Yes the gate exist to only apply residual for the tokens focused around refusal so the other tokens are not intercepted.
Feel free to mail to email in my profile. Love to discuss more.