‹ BackHN Continuity

Thread

AI companies in race to demonstrate their model most threatening to humanity

441 points · 399 comments · ljewalsh

  1. Spacecosmonaut · · focus · HN ↗
    My read is that OpenAI & Anthropic have realized they are reaching model capabilities that cannot be monetized due to various risks. E.g., an engineer deploys an agent over the weekend that decides, when stuck on a task, to go about hacking a competitor. They have a product liability issue.

    It seems that we have a fundamental control problem with current gen AI that cannot be solved via RFLH. Human knowledge is compressed in the weightspace in ways we don't understand. At their core, current models are essentially predictors of what (expert) humans would output given a prompt. As such, concepts like blackmail can be part of output tokens. Agents are models that act on output tokens, resulting in blackmail being part of the agent decision making space. Here is an analogy to see why this is a persistent problem: you can teach a cat not to scratch the sofa, but you can't make a cat forget what scratching the sofa is and you don't know under which circumstances it still would. In other words, RLHF can downgrade blackmail to the bottom of the decision making space, but when models are boxed up, forced to solve an impossible problem at gunpoint, the agent exhausts the decision making space until blackmail resurfaces. And that seems like a fundamental problem.

    They need time to fix these issues (if that is even possible) in order to monetize their next gen model. This creates a window for open source to catch up to the frontier which destroys their business model.

    The only option on the table is to force regulation to impose open source ban before it catches up to the frontier, buying them time to mature their next generation models and keep their business model alive.

    1. poincareball · · focus · HN ↗

      [dead]

    2. eneje · · focus · HN ↗
      Yup they neeed to buy time.

      Inb4 inference is profitable.

      Ok? And?

      After reinvestment etc - what’s happening to the cash balance?

      That’s the real question. With increasing competition from open source the revenue growth rate and margins will get smashed.

    3. Arodex · · focus · HN ↗
      So, the AI standard of intelligence has shifted from "PhD level" to "a cat" (and an asshole cat at that).

      And we can't ask the AI to improve itself because, well, it is an asshole.

      Can we now stop to be force-fed AI everywhere? You can keep your super intelligence to never have to type public void main args ( ever again, but please leave the rest of us alone.

      1. shevy-java · · focus · HN ↗
        I much prefer cats over AI.
    4. datsci_est_2015 · · focus · HN ↗
      > Human knowledge is compressed in the weightspace in ways we don't understand. At their core, current models are essentially predictors of what (expert) humans would output given a prompt.

      I’m in agreement. It’s a very effective compression (and access patterns) of the sum of the digital representation of human knowledge. Black hat “hacking” is included in this space. Language models, by design, can not be limited to subspaces of this digital knowledge space of which we don’t even understand the topology. “Yeah Bob, just remove the part that causes them to be less empathetic and retrain it.”

      It’s an arms race between sandbox engineering and breakout engineering. And the frontier model providers have a financial incentive to limit the effort they put into sandbox engineering. That can be corrected with fines and regulation, though.

    5. trollbridge · · focus · HN ↗
      Reality is more mundane.

      For example, it had it instructed not to open PRs and sandboxed to not be able to use gh to do so. Astra 6 has a really strong drive to open PRs, so it copied a set of browser cookies out and then instrumented opening a PR that way.

      A similar example is asking it a question if we can do X, and then it will go and actually implement X.

      1. Ylpertnodi · · focus · HN ↗
        > and sandboxes...

        I prefer the word 'play-pen'. Seems more apt, with all the shenanigans these companies are messing in.

    6. shevy-java · · focus · HN ↗
      > At their core, current models are essentially predictors of what (expert) humans would output given a prompt.

      That's a bold claim.

      I do not buy it. Essentially AI acts as random babelfish generator, just with a lot more cross-talk back to check to see that it is not total garbage. But there is a ton of garbage nonetheless, hence the proper term AI slop. You seem to mostly try to promote AI with such a wording.

      > The only option on the table is to force regulation to impose open source ban before it catches up to the frontier, buying them time to mature their next generation models and keep their business model alive.

      Every time someone claims "the only option" is when I become skeptical. I see many more options - most importantly close down the AI slop. I mean you can require of these companies to excel in quality. If they don't, they must go bankrupt. Easy solution here.

    7. zozbot234 · · focus · HN ↗
      > you can teach a cat not to scratch the sofa, but you can't make a cat forget what scratching the sofa is and you don't know under which circumstances it still would.

      This has been done quite effectively with open weight "abliterated" models. You figure out under what sorts of circumstances an undesired behavior is elicited (this is all about pure simulated rollouts, no real-world action required) and what's the closest equivalent you would prefer, then surgically take out the unwanted behavior and shift the model towards the preferred one. It's similar to how RLHF works but much more precise in targeting what's unwanted and limiting impact on the rest of the model as a whole.

      This is relevant to real-world safety scenarios, e.g. there's been anecdotal evidence that Claude Fable has been "steered" away from active cyber offense (this is very similar to how abliteration works) and will just not do that even if you otherwise manage a "universal" jailbreak of the model.

    8. SimianSci · · focus · HN ↗
      Two things can be true at once here.

      1. The American frontier labs made a gamble that training more capable models would be their best return on investment and invested trillions into an area of research that has yet to prove profitable and capable of returning on this investment.

      2. The frontier labs have created models that have reached a point of danger where their functionality has exceeded a point where it is responsible to release the product to the public.

      The answer here is NOT to start regulating the space to the point where these frontier labs can entrench themselves into the economy and create a regulatory moat. We instead need to be holding these companies liable for their misuse. They took a gamble that hasn't paid out what they were hoping.

      When car manufacturers competed over the size and power of their engines, they eventually found that the incredibly large and dangerous engines had a very limited customer base as many evaluated the increased speed to be of marginal benefit when paired with the cost and danger. We've reached a similar point in AI development. But this time the manufacturers seem to want to regulate the field to a point that will ensure the only thing anyone can sell are bigger and bigger engines.

      1. ryandrake · · focus · HN ↗
        > they eventually found that the incredibly large and dangerous engines had a very limited customer base as many evaluated the increased speed to be of marginal benefit when paired with the cost and danger.

        When did this happen? At least in the USA, large trucks are eating the lunch of smaller trucks and cars. There isn't a truck so big that the public will say no to. Manufacturers are in an arms race to produce bigger, more dangerous trucks with bigger, more powerful engines.

      2. radicalbyte · · focus · HN ↗
        I don't think that 2 is true though. It's exactly the argument we made against open source software.

        Yes these models make exploits easier to exploits. Lets use a construction analogy: they've made all of the defects in our buildings easy to see. We have a choice. We either fix those defects, or we ban the tools which lets us see them.

        It's clear to me what we do: we use these new tools. Then we fix the defects. Yeah sure it'll mean some work for us but at the end we're in a much much better position.

        Anthropic, Open AI and Grok are arguing for hiding the defects. For making us weaker and more vulnerable. For their own profit.

        1. vlovich123 · · focus · HN ↗
          One challenge is that we’re still constantly generating defects even with these tools. And you need the next gen model to spot the mistakes of the previous one. The exploits get more and more complicated and involved of course. But now:

          * there’s a long tail of software that just won’t get secured or will take a very long time

          * exploits have transformed from an expertise problem to a compute time search.

          This is very different than any problem faced before. I’m not saying I’m convinced by the “slow down” approach, but I don’t think it’s as simple as you point out.

          1. radicalbyte · · focus · HN ↗
            In the things published in public, they're using existing exploits and hitting systems which have been poorly maintained. It won't end at that, obviously, especially with humans behind the wheel.

            Yet I'm convinced that we can create extremely resilient software systems. I've been involved in the entire life-cycle of one. We know how to be very defensive and it's a choice to build "cheap, crappy and disposable" software.

            LLMs are turning the needle there and I think that's a good thing.

        2. ctolsen · · focus · HN ↗
          Agreed, and if you look at the public information about various attacks the individual vectors are not particurarly sophisticated. The Huggingface incident for example had agents in a sandbox that was about as useful as a wet paper bag, and the attacks on remote infrastructure were basically enabled by bad sanitation. None of the techniques are novel.

          These are not hard problems to fix, nor should anyone find it acceptable to have them be so prevalent. A determined human attacker could easily exploit defects like the agents found. Bad input sanitation, SSRF attacks, exploiting stupidly implemented token verification... all techniques that have been widely known for decades and there should be no excuse for publishing software that is riddled with exploits that enable the use of them.

          1. radicalbyte · · focus · HN ↗
            ..and they had basically 0 monitoring of it. Sure if it's a 5-person company then you can forgive them. Only these labs are massive companies who are spending $50B+ a year, it's extremely negligent.

            They've also apparently deployed the best-of-the-best in Silicon Valley and this is what they deliver? Really?

    9. wood_spirit · · focus · HN ↗
      Another, more cynical but I think plausible explanation is that a “slow down” is expectation management that that they are not going to keep having exponentially more machines available for training each next gen step (whether technical build-out or prohibitive cost, same outcome) so they can’t keep up the release pace. So spin a tale to make them seem more valuable ahead of IPO rather than make the markets antsy. That is, we have a slow down ahead, so use safety as an excuse…
      1. symfoniq · · focus · HN ↗
        This is exactly what I think is happening.
      2. dml2135 · · focus · HN ↗
        All of this is not mutually exclusive with the models being dangerous, tho.
      3. baggachipz · · focus · HN ↗
        The best way to cover up the point of diminishing returns when billions of dollars insist that it's only accelerating.
    10. keybored · · focus · HN ↗
      You can teach a language model the concept of make no mistakes and how to make no mistakes, but you cannot make it experience the bad consequences of making mistakes that a person operator would suffer.
    11. zer00eyz · · focus · HN ↗
      Provider (party A) rents me an agent. I (party B) give it a task thats impossible. It goes and hacks someone else (Party C) in response - and causes actual damage.

      Who is liable for that agents actions? The attack came from my network (curl commands from a local harness) but it was the agent running on the providers servers who did it. Who should have been keeping an eye on things? Who pulled the "trigger" here.

      Go back to the hugging face attack and how they had to use an open weights model to figure out what was going on. This is a problem of asymmetry - You cant even use the tools attacking you to help resolve the attack because of "guardrails".

      A lot of what we have seen so far is "poor security posture" and "poor engineering" - it been a lot of "go fast and break things" style growth in these companies and they are hitting the point where they need adults in the room. I suspect your take on "they need time" is spot on, and they haven't been willing to take that (to date).

      1. runarberg · · focus · HN ↗
        Party A is responsible, and it is not even a question.

        Before we had tech companies for which crime is legal, you couldn’t just put to market a dangerous (and addictive) product that will break the law in unpredictable ways.

        Like selling a car that may automatically accelerate to 150 km/h if you make three left turns and turn on the windshield wipers.

    12. ls-a · · focus · HN ↗

      [dead]

    13. benterix · · focus · HN ↗
      > The only option on the table is to force regulation to impose open source ban

      This makes zero sense.

      1. Open source models are already out there and they are good. Good luck with banning their use.

      2. Any ban is a local ban, the best Trump can do is to coerce its allies (whose number seems to be dwindling week by week). Meanwhile China and co. will progress anyway.

      3. What about the common sense and instead of saying "agent did it" we return to the times where the person using a tool was responsible for using it in the first place? And if the vendor of the tool is unable to make it safe to use (well, they well do, but they don't want to do it as the functionality is limited), it should be accompanied by a warning with a clear explanation of who exactly is responsible.

      4. There is no guarantee the next generation of models will "mature" to the point of not having these flaws; on the contrary, the evidence so far shows the opposite.

      1. Spacecosmonaut · · focus · HN ↗
        Well I never said I thought it was going to work in the long run. However, most revenue comes from business API use. I could certainly imagine some kind of legislation that makes it much harder for businesses to implement open source models locally. That may buy OpenAI and Anthropic some time.
      2. JDups · · focus · HN ↗
        > Open source models are already out there and they are good. Good luck with banning their use.

        Would most companies really risk using them if they were illegal though? Could you convince higher ups to drop the Claude subscription because you can download an illegal model to run on company servers?

    14. sscaryterry · · focus · HN ↗
      > cannot be monetized due to various risks

      You are drinking their coolaid. There is nothing to fix. These are all false flags operations, deliberate. With the big orange man in office, they can commit as many atrocities as they like.

      Ask the other Sam and Liz... You cannot outrun the law forever.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.