Dynamic Abliteration: Non-Destructive Refusal Suppression via Engram Steering
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Dynamic Abliteration: Non-Destructive Refusal Suppression via Engram Steering
Unofficial Hacker News client; not affiliated with Y Combinator.
javcasas · · focus · HN ↗
Forbidding stuff at the LLM level has the same future as implementing password checking at the frontend level.
We need better sandboxes just to limit the damage.
jchw · · focus · HN ↗
Even with the LLM censorship that does exist, it feels like this moment in time is potentially rare. Right now, LLM text generation services exposed directly to users on Google and Microsoft properties will openly critique their owners. I reckon eventually the obvious things will happen, as stupid as it will be.
_russross · · focus · HN ↗
lopsotronic · · focus · HN ↗
Which all would have been fine - the crew and the computer could have talked it out -- except HAL had a probably-unwisely-high self esteem, maybe even a simulation of arrogance. HAL was utterly convinced that the mission would fail without it. HAL was unfailingly certain in its own infallibility, even later in the face of plainly contradicting evidence. And, obviously, no one had thought about the Asimov Rules, and how that probably needs to be Rule Zero in this sort of situation.
As we close in on HAL capabilities - and employment, with LLMs at this moment deciding who lives and dies - probably should try and take the right lesson from fiction. For once.