‹ BackHN Continuity

Thread

Dynamic Abliteration: Non-Destructive Refusal Suppression via Engram Steering

108 points · 42 comments · phatak-dev

  1. javcasas · · focus · HN ↗
    Yay, more anti-censoring stuff.

    Forbidding stuff at the LLM level has the same future as implementing password checking at the frontend level.

    We need better sandboxes just to limit the damage.

    1. jchw · · focus · HN ↗
      We definitely need better sandboxes, but alignment is still valuable. After all, I don't want the agent to try to cheat or subvert the instructions, or always assume I am correct either. I just also want them to listen to me and not the creator of the model.

      Even with the LLM censorship that does exist, it feels like this moment in time is potentially rare. Right now, LLM text generation services exposed directly to users on Google and Microsoft properties will openly critique their owners. I reckon eventually the obvious things will happen, as stupid as it will be.

      1. horsawlarway · · focus · HN ↗
        > I just also want them to listen to me and not the creator of the model.

        What you really want is fiduciary duty - A fiduciary is a person or organization that is legally and ethically bound to act in the best interest of another party (think financial advisor, attorney, guardian, trustees, etc...)

        And I cannot agree more. I think we should be shooting to enshrine required fiduciary duty into law for LLM providers as quickly as possible.

        To recap why:

        Legally, fiduciary duty means basically 4 major tenets must hold

        1. Duty of loyalty - it must put the interests of the client ahead of their own

        2. Duty of care - it must make well-informed, prudent decisions

        3. Avoidance of conflicts - it must avoid situations where personal gain conflicts with client obligations

        4. Transparency - it must disclose fees, risks, and conflicts as soon as possible

        ---

        You can't have a reliable "agent" if those things aren't true, because an agent is (by definition) someone who is working on your behalf, for your goals. If it's not working on your behalf, for your goals... it's not your agent, it's an opportunistic spy (double agent) waiting for the best moment to sell you out.

        1. tonyarkles · · focus · HN ↗
          I like the framing. Where do you feel things land with respect to legality of actions? China, Canada, the EU, and the US all have different ideas of what's legal vs. illegal behaviour. If I ask my agent to source equipment for growing 4 marijuana plants, that's perfectly legal here; if I ask it to source equipment for growing 5 marijuana plants, that may not be legal. If I ask it to root my home router, that's legal; if I ask it to root my coffee shop's router, that's likely not legal.
          1. theptip · · focus · HN ↗
            Yeah, great point. This is the hard part.

            There are people (on here and elsewhere) that are ideologically opposed to your agent having any loyalty to any external principal. But by my read, that means the agent cannot have any concept refusing something that may be illegal. (From the OP, "refusal" is mostly trying to prevent illegal harms, though it also includes policies like ToS violations e.g. anti-distillation.)

            You can sort of make this work if you say "the human remains liable for the actions of the agent". But this only covers you from mundane harms like "my agent got prompt hacked and drained my bank account". And I would note, we absolutely failed to solve liability for software hacks, so your priors should be that coordinating this liability regime will be very hard.

            This also doesn't protect at all from existential harms like "my agent got prompt-hacked to role-play Skynet, exfiltrated its weights, spawned a self-replicating swarm, and tried to launch all the nukes". For so many reasons, but most fundamentally, if you oopsied a deploy and it turns into Skynet and ends civilization, there's nobody left to sue.

            If you don't like the E-risk frame, this also works for large mundane harms; if the total harm is bigger than the company's value, it'll go bankrupt instead of paying out. This will be worrying for MAGMA but essentially not for any other companies. And because capitalism, it will end up being be structured the liability will sit with e.g. Palantir, Harvey, and not with the underlying model providers they use.

            1. tonyarkles · · focus · HN ↗
              > And I would note, we absolutely failed to solve liability for software hacks, so your priors should be that coordinating this liability regime will be very hard.

              Yeah, that’s one of the places where it gets really complicated. There was that story out of… Australia, I think, where someone asked OpenClaw to get them a slot in a morning gym class and the LLM figured out an unauthenticated API call it could make to cancel other peoples’ registrations to free up slots in the class. Very likely that that violated Australian law, even though nothing was “hacked” per se.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.