‹ BackHN Continuity

Thread

AI companies in race to demonstrate their model most threatening to humanity

441 points · 399 comments · ljewalsh

  1. Spacecosmonaut · · focus · HN ↗
    My read is that OpenAI & Anthropic have realized they are reaching model capabilities that cannot be monetized due to various risks. E.g., an engineer deploys an agent over the weekend that decides, when stuck on a task, to go about hacking a competitor. They have a product liability issue.

    It seems that we have a fundamental control problem with current gen AI that cannot be solved via RFLH. Human knowledge is compressed in the weightspace in ways we don't understand. At their core, current models are essentially predictors of what (expert) humans would output given a prompt. As such, concepts like blackmail can be part of output tokens. Agents are models that act on output tokens, resulting in blackmail being part of the agent decision making space. Here is an analogy to see why this is a persistent problem: you can teach a cat not to scratch the sofa, but you can't make a cat forget what scratching the sofa is and you don't know under which circumstances it still would. In other words, RLHF can downgrade blackmail to the bottom of the decision making space, but when models are boxed up, forced to solve an impossible problem at gunpoint, the agent exhausts the decision making space until blackmail resurfaces. And that seems like a fundamental problem.

    They need time to fix these issues (if that is even possible) in order to monetize their next gen model. This creates a window for open source to catch up to the frontier which destroys their business model.

    The only option on the table is to force regulation to impose open source ban before it catches up to the frontier, buying them time to mature their next generation models and keep their business model alive.

    1. zozbot234 · · focus · HN ↗
      > you can teach a cat not to scratch the sofa, but you can't make a cat forget what scratching the sofa is and you don't know under which circumstances it still would.

      This has been done quite effectively with open weight "abliterated" models. You figure out under what sorts of circumstances an undesired behavior is elicited (this is all about pure simulated rollouts, no real-world action required) and what's the closest equivalent you would prefer, then surgically take out the unwanted behavior and shift the model towards the preferred one. It's similar to how RLHF works but much more precise in targeting what's unwanted and limiting impact on the rest of the model as a whole.

      This is relevant to real-world safety scenarios, e.g. there's been anecdotal evidence that Claude Fable has been "steered" away from active cyber offense (this is very similar to how abliteration works) and will just not do that even if you otherwise manage a "universal" jailbreak of the model.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.