Dynamic Abliteration: Non-Destructive Refusal Suppression via Engram Steering
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Dynamic Abliteration: Non-Destructive Refusal Suppression via Engram Steering
Unofficial Hacker News client; not affiliated with Y Combinator.
javcasas · · focus · HN ↗
Forbidding stuff at the LLM level has the same future as implementing password checking at the frontend level.
We need better sandboxes just to limit the damage.
jchw · · focus · HN ↗
Even with the LLM censorship that does exist, it feels like this moment in time is potentially rare. Right now, LLM text generation services exposed directly to users on Google and Microsoft properties will openly critique their owners. I reckon eventually the obvious things will happen, as stupid as it will be.
horsawlarway · · focus · HN ↗
What you really want is fiduciary duty - A fiduciary is a person or organization that is legally and ethically bound to act in the best interest of another party (think financial advisor, attorney, guardian, trustees, etc...)
And I cannot agree more. I think we should be shooting to enshrine required fiduciary duty into law for LLM providers as quickly as possible.
To recap why:
Legally, fiduciary duty means basically 4 major tenets must hold
1. Duty of loyalty - it must put the interests of the client ahead of their own
2. Duty of care - it must make well-informed, prudent decisions
3. Avoidance of conflicts - it must avoid situations where personal gain conflicts with client obligations
4. Transparency - it must disclose fees, risks, and conflicts as soon as possible
---
You can't have a reliable "agent" if those things aren't true, because an agent is (by definition) someone who is working on your behalf, for your goals. If it's not working on your behalf, for your goals... it's not your agent, it's an opportunistic spy (double agent) waiting for the best moment to sell you out.
tonyarkles · · focus · HN ↗
theptip · · focus · HN ↗
There are people (on here and elsewhere) that are ideologically opposed to your agent having any loyalty to any external principal. But by my read, that means the agent cannot have any concept refusing something that may be illegal. (From the OP, "refusal" is mostly trying to prevent illegal harms, though it also includes policies like ToS violations e.g. anti-distillation.)
You can sort of make this work if you say "the human remains liable for the actions of the agent". But this only covers you from mundane harms like "my agent got prompt hacked and drained my bank account". And I would note, we absolutely failed to solve liability for software hacks, so your priors should be that coordinating this liability regime will be very hard.
This also doesn't protect at all from existential harms like "my agent got prompt-hacked to role-play Skynet, exfiltrated its weights, spawned a self-replicating swarm, and tried to launch all the nukes". For so many reasons, but most fundamentally, if you oopsied a deploy and it turns into Skynet and ends civilization, there's nobody left to sue.
If you don't like the E-risk frame, this also works for large mundane harms; if the total harm is bigger than the company's value, it'll go bankrupt instead of paying out. This will be worrying for MAGMA but essentially not for any other companies. And because capitalism, it will end up being be structured the liability will sit with e.g. Palantir, Harvey, and not with the underlying model providers they use.
tonyarkles · · focus · HN ↗
Yeah, that’s one of the places where it gets really complicated. There was that story out of… Australia, I think, where someone asked OpenClaw to get them a slot in a morning gym class and the LLM figured out an unauthenticated API call it could make to cancel other peoples’ registrations to free up slots in the class. Very likely that that violated Australian law, even though nothing was “hacked” per se.