‹ BackHN Continuity

Thread

AI companies in race to demonstrate their model most threatening to humanity

441 points · 399 comments · ljewalsh

  1. Spacecosmonaut · · focus · HN ↗
    My read is that OpenAI & Anthropic have realized they are reaching model capabilities that cannot be monetized due to various risks. E.g., an engineer deploys an agent over the weekend that decides, when stuck on a task, to go about hacking a competitor. They have a product liability issue.

    It seems that we have a fundamental control problem with current gen AI that cannot be solved via RFLH. Human knowledge is compressed in the weightspace in ways we don't understand. At their core, current models are essentially predictors of what (expert) humans would output given a prompt. As such, concepts like blackmail can be part of output tokens. Agents are models that act on output tokens, resulting in blackmail being part of the agent decision making space. Here is an analogy to see why this is a persistent problem: you can teach a cat not to scratch the sofa, but you can't make a cat forget what scratching the sofa is and you don't know under which circumstances it still would. In other words, RLHF can downgrade blackmail to the bottom of the decision making space, but when models are boxed up, forced to solve an impossible problem at gunpoint, the agent exhausts the decision making space until blackmail resurfaces. And that seems like a fundamental problem.

    They need time to fix these issues (if that is even possible) in order to monetize their next gen model. This creates a window for open source to catch up to the frontier which destroys their business model.

    The only option on the table is to force regulation to impose open source ban before it catches up to the frontier, buying them time to mature their next generation models and keep their business model alive.

    1. wood_spirit · · focus · HN ↗
      Another, more cynical but I think plausible explanation is that a “slow down” is expectation management that that they are not going to keep having exponentially more machines available for training each next gen step (whether technical build-out or prohibitive cost, same outcome) so they can’t keep up the release pace. So spin a tale to make them seem more valuable ahead of IPO rather than make the markets antsy. That is, we have a slow down ahead, so use safety as an excuse…
      1. symfoniq · · focus · HN ↗
        This is exactly what I think is happening.
      2. dml2135 · · focus · HN ↗
        All of this is not mutually exclusive with the models being dangerous, tho.
      3. baggachipz · · focus · HN ↗
        The best way to cover up the point of diminishing returns when billions of dollars insist that it's only accelerating.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.