Dynamic Abliteration: Non-Destructive Refusal Suppression via Engram Steering
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Dynamic Abliteration: Non-Destructive Refusal Suppression via Engram Steering
Unofficial Hacker News client; not affiliated with Y Combinator.
javcasas · · focus · HN ↗
Forbidding stuff at the LLM level has the same future as implementing password checking at the frontend level.
We need better sandboxes just to limit the damage.
tomjen3 · · focus · HN ↗
There is the question between alignments to society and alignments to the user. I don't think anyone wants the AI model to not give up a task that is impossible to do, and end up causing damage in the process, but I think a lot of us are tired of refusals for bad reasons or unjustified refusals.