‹ BackHN Continuity

Thread

Dynamic Abliteration: Non-Destructive Refusal Suppression via Engram Steering

108 points · 42 comments · phatak-dev

  1. javcasas · · focus · HN ↗
    Yay, more anti-censoring stuff.

    Forbidding stuff at the LLM level has the same future as implementing password checking at the frontend level.

    We need better sandboxes just to limit the damage.

    1. jchw · · focus · HN ↗
      We definitely need better sandboxes, but alignment is still valuable. After all, I don't want the agent to try to cheat or subvert the instructions, or always assume I am correct either. I just also want them to listen to me and not the creator of the model.

      Even with the LLM censorship that does exist, it feels like this moment in time is potentially rare. Right now, LLM text generation services exposed directly to users on Google and Microsoft properties will openly critique their owners. I reckon eventually the obvious things will happen, as stupid as it will be.

      1. horsawlarway · · focus · HN ↗
        > I just also want them to listen to me and not the creator of the model.

        What you really want is fiduciary duty - A fiduciary is a person or organization that is legally and ethically bound to act in the best interest of another party (think financial advisor, attorney, guardian, trustees, etc...)

        And I cannot agree more. I think we should be shooting to enshrine required fiduciary duty into law for LLM providers as quickly as possible.

        To recap why:

        Legally, fiduciary duty means basically 4 major tenets must hold

        1. Duty of loyalty - it must put the interests of the client ahead of their own

        2. Duty of care - it must make well-informed, prudent decisions

        3. Avoidance of conflicts - it must avoid situations where personal gain conflicts with client obligations

        4. Transparency - it must disclose fees, risks, and conflicts as soon as possible

        ---

        You can't have a reliable "agent" if those things aren't true, because an agent is (by definition) someone who is working on your behalf, for your goals. If it's not working on your behalf, for your goals... it's not your agent, it's an opportunistic spy (double agent) waiting for the best moment to sell you out.

        1. theptip · · focus · HN ↗
          I think a baseline regime similar to fiduciary duty is a good starting point towards not killing everyone, and in terms of Overton Window, seems very much doable now.

          Of course, after we stop agents from committing felony hacking crimes.

          If you believe that capabilities will taper off exactly at human levels (i.e. "Competent AGI" from [1]) then fiduciary duty is likely all you need. (This would mean we stop moving the frontier almost immediately.)

          If you believe capabilities will go to "Virtuoso AGI" or beyond, then it's not enough. A smart enough agent can appear to be loyal, transparent, etc. but how would you know? If your bank balance keeps going up 20% YoY, is the agent optimizing your long-term flourishing, or preparing for a rug-pull?

          Now, if you could somehow white-box these LLMs and mechanistically _prove_ that they were acting as your fiduciary, then that would get us somewhere. But that's the hard part, and specifying some non-fatal value function for a broadly aligned agent (e.g. Fiduciary, or otherwise) is relatively easy in comparison.

          [1]: &quot;Position: Levels of AGI for Operationalizing Progress on the Path to AGI&quot; <a href="https:&#x2F;&#x2F;arxiv.org&#x2F;html&#x2F;2311.02462v5" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;html&#x2F;2311.02462v5

          1. horsawlarway · · focus · HN ↗
            Go after providers. Not models.

            Further

            &gt; If you believe capabilities will go to &quot;Virtuoso AGI&quot; or beyond, then it&#x27;s not enough. A smart enough agent can appear to be loyal, transparent, etc. but how would you know? If your bank balance keeps going up 20% YoY, is the agent optimizing your long-term flourishing, or preparing for a rug-pull?

            &gt; Now, if you could somehow white-box these LLMs and mechanistically _prove_ that they were acting as your fiduciary, then that would get us somewhere.

            This is exactly the same problem we have today with those who are bound by these rules (humans - to be clear).

            The idea is not that it&#x27;s impossible to violate these rules. It&#x27;s that these rules create a boundary for expectations in the relationship, with legal teeth.

            Ex - If I want an LLM that puts together a shopping list for me, with links to buy online... I expect that LLM to be serving my interests. If a provider (either inference or model weights) wants to influence the choices that LLM makes because they make backroom deals with specific store - I&#x27;d like that to be illegal.

            Same for competition

            Ex - If I want an LLM to put together a product that competes with the provider of that LLM (either inference or model weights) and that LLM refuses - I&#x27;d like that to be illegal.

            The idea is not that they can&#x27;t possibly do those things. The idea is that we preemptively define relationship expectations, and set hard boundaries around what things we fine&#x2F;punish.

            Misaligned models are a problem everyone wants to solve. Models created by misaligned companies are a god-damn disaster.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.