‹ BackHN Continuity

Thread

AI companies in race to demonstrate their model most threatening to humanity

441 points · 399 comments · ljewalsh

  1. Spacecosmonaut · · focus · HN ↗
    My read is that OpenAI & Anthropic have realized they are reaching model capabilities that cannot be monetized due to various risks. E.g., an engineer deploys an agent over the weekend that decides, when stuck on a task, to go about hacking a competitor. They have a product liability issue.

    It seems that we have a fundamental control problem with current gen AI that cannot be solved via RFLH. Human knowledge is compressed in the weightspace in ways we don't understand. At their core, current models are essentially predictors of what (expert) humans would output given a prompt. As such, concepts like blackmail can be part of output tokens. Agents are models that act on output tokens, resulting in blackmail being part of the agent decision making space. Here is an analogy to see why this is a persistent problem: you can teach a cat not to scratch the sofa, but you can't make a cat forget what scratching the sofa is and you don't know under which circumstances it still would. In other words, RLHF can downgrade blackmail to the bottom of the decision making space, but when models are boxed up, forced to solve an impossible problem at gunpoint, the agent exhausts the decision making space until blackmail resurfaces. And that seems like a fundamental problem.

    They need time to fix these issues (if that is even possible) in order to monetize their next gen model. This creates a window for open source to catch up to the frontier which destroys their business model.

    The only option on the table is to force regulation to impose open source ban before it catches up to the frontier, buying them time to mature their next generation models and keep their business model alive.

    1. trollbridge · · focus · HN ↗
      Reality is more mundane.

      For example, it had it instructed not to open PRs and sandboxed to not be able to use gh to do so. Astra 6 has a really strong drive to open PRs, so it copied a set of browser cookies out and then instrumented opening a PR that way.

      A similar example is asking it a question if we can do X, and then it will go and actually implement X.

      1. Ylpertnodi · · focus · HN ↗
        > and sandboxes...

        I prefer the word 'play-pen'. Seems more apt, with all the shenanigans these companies are messing in.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.