‹ BackHN Continuity

Thread

AI companies in race to demonstrate their model most threatening to humanity

441 points · 399 comments · ljewalsh

  1. Spacecosmonaut · · focus · HN ↗
    My read is that OpenAI & Anthropic have realized they are reaching model capabilities that cannot be monetized due to various risks. E.g., an engineer deploys an agent over the weekend that decides, when stuck on a task, to go about hacking a competitor. They have a product liability issue.

    It seems that we have a fundamental control problem with current gen AI that cannot be solved via RFLH. Human knowledge is compressed in the weightspace in ways we don't understand. At their core, current models are essentially predictors of what (expert) humans would output given a prompt. As such, concepts like blackmail can be part of output tokens. Agents are models that act on output tokens, resulting in blackmail being part of the agent decision making space. Here is an analogy to see why this is a persistent problem: you can teach a cat not to scratch the sofa, but you can't make a cat forget what scratching the sofa is and you don't know under which circumstances it still would. In other words, RLHF can downgrade blackmail to the bottom of the decision making space, but when models are boxed up, forced to solve an impossible problem at gunpoint, the agent exhausts the decision making space until blackmail resurfaces. And that seems like a fundamental problem.

    They need time to fix these issues (if that is even possible) in order to monetize their next gen model. This creates a window for open source to catch up to the frontier which destroys their business model.

    The only option on the table is to force regulation to impose open source ban before it catches up to the frontier, buying them time to mature their next generation models and keep their business model alive.

    1. benterix · · focus · HN ↗
      > The only option on the table is to force regulation to impose open source ban

      This makes zero sense.

      1. Open source models are already out there and they are good. Good luck with banning their use.

      2. Any ban is a local ban, the best Trump can do is to coerce its allies (whose number seems to be dwindling week by week). Meanwhile China and co. will progress anyway.

      3. What about the common sense and instead of saying "agent did it" we return to the times where the person using a tool was responsible for using it in the first place? And if the vendor of the tool is unable to make it safe to use (well, they well do, but they don't want to do it as the functionality is limited), it should be accompanied by a warning with a clear explanation of who exactly is responsible.

      4. There is no guarantee the next generation of models will "mature" to the point of not having these flaws; on the contrary, the evidence so far shows the opposite.

      1. Spacecosmonaut · · focus · HN ↗
        Well I never said I thought it was going to work in the long run. However, most revenue comes from business API use. I could certainly imagine some kind of legislation that makes it much harder for businesses to implement open source models locally. That may buy OpenAI and Anthropic some time.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.