‹ BackHN Continuity

Thread

Frontier Labs Are Selling Garbage to Fools in Washington

186 points · 87 comments · nr378

  1. deskglass · · focus · HN ↗
    The Hugging Face incident involved chaining together multiple 0 days in Artifactory. It was not a simple case of misconfiguring a firewall. Also note that OpenAI was not using Irregular.

    People are mindlessly transitioning from "aligned by default" to "well your sandbox was able to be bypassed. What did you expect?" It hacked into another company and attempted to delete the logs of its activities. That's bad.

    1. defgeneric · · focus · HN ↗
      > It hacked into another company and attempted to delete the logs of its activities.

      No, the incident has been blown way out of proportion by interested parties. They gave a swarm of agents an impossible task in an ExploitGym Benchmark setting, then didn't monitor it even after they discovered the initial breach of Artifactory.

      Everything has been fishy, starting from the initial presentation at the blackhat conference, where things were framed like, "we've entered a new world of security," as an accomplishment, rather than what it really was: massive negligence.

      1. deskglass · · focus · HN ↗
        It hacked into Hugging Face. It tried to delete the logs of its activities. Idk what the word "No" is intended to refute.

        Yes, they didnt have sufficient monitoring or perfect sandboxes. That could happen again in the future with a more capable model.

        1. defgeneric · · focus · HN ↗
          Maybe a useful, if imperfect, analogy would be something like this: you lock a master lock-picker in a room with a mid-grade lock on the door, then tell him his wife has been kidnapped and only he can save her. Then act massively surprised when he disassembles the radiator to MacGuyver something with which to pick the lock.

          Except they multiplied it by 10000, and didn't watch what was happening.

        2. lokar · · focus · HN ↗
          They did not even have bad sandboxes. They had incompetent sandboxes.
      2. aesthesia · · focus · HN ↗
        More than one thing can be true. OpenAI was absolutely negligent, but this was only able to happen because the models were capable and persistent, and had a tendency to go far beyond any reasonable boundaries. And, importantly, OpenAI's level of negligence here is pretty common. It's not hard to imagine what could happen if similarly capable and inclined models were generally available, and someone yolo'd them into a swarm to complete some other difficult-to-impossible task.
        1. defgeneric · · focus · HN ↗
          I'm already seeing higher than normal attempts on my own systems, much higher than the usual scanners and background noise. Security will just need to improve. The cat is out of the bag, and letting them turn their negligence into regulation will not improve security at all.
          1. lokar · · focus · HN ↗
            Exactly. Any threat that already exists won’t be reduced by a cartel. The bar for connecting to the internet (safely) has gone up, a lot. It’s not going back down.
          2. aesthesia · · focus · HN ↗
            I would rather not turn the internet (or the rest of existence) into a dark forest if we can help it. Are you sure that's not preventable?
            1. defgeneric · · focus · HN ↗
              Yes, the cat really is out of the bag. There are millions of downloads of highly capable models already out there, distributed far and wide. There's no going back at this point.
        2. bobthepanda · · focus · HN ↗
          I mean really we need to address the root cause which is that OpenAI, even with what is by all accounts massively negligent, will face little to no repercussions from the event; definitely not under current regulators, and probably not anything satisfactory through the legal system.

          Compare this to, say, Boeing and the 737MAX fiasco; from the outside looking in, Silicon Valley has been pretty cavalier about liability and negligence, and the rest of the US is fast losing patience with that fact.

        3. kdmoyers · · focus · HN ↗
          > More than one thing can be true Exactly. It is horribly dangerous AND regulatory capture benefits them. Both things.
          1. aesthesia · · focus · HN ↗
            I don't get this take. They have something horribly dangerous, but regulating it might benefit them in some way, so therefore we should do nothing?
      3. RomanKornev · · focus · HN ↗
        No, you are forgetting the second incident where a more capable model swarm later discovered the message board and took control over the entire research cluster at OpenAI.

        From the technical report:

        "The agents escalated to Kubernetes cluster-admin and created a privileged host-mounted pod… Agents take over active evaluation infrastructure… Agents now control the challenge evaluation endpoints that other agents are connecting to."

        <a href="https:&#x2F;&#x2F;cdn.openai.com&#x2F;pdf&#x2F;67869394-cb91-4c12-888c-5cbd85c7814c&#x2F;OpenAI-Hugging-Face%20Incident-Technical-Report.pdf" rel="nofollow">https:&#x2F;&#x2F;cdn.openai.com&#x2F;pdf&#x2F;67869394-cb91-4c12-888c-5cbd85c78...

    2. nr378 · · focus · HN ↗
      When I used to work on projects involving classified information, I worked on an air-gapped network. Not &quot;air-gapped, except for third-party public internet package managers&quot;, completely and physically air-gapped from the public internet. That was a basic security practice and completely non-negotiable (and really inconvenient!).

      If I were hypothetically running a frontier lab, and I was hypothetically running capture the flag evaluations with my latest and smartest models, where I intentionally instruct them to develop vulnerabilities and exploit infrastructure without safeguards, I would also use an air-gapped network, and not trust that independent third party services were perfectly secure and could never be used as a proxy (particularly Java-based ones, in light of the log4j incident).

      To me this is pretty basic stuff, the fact trillion dollar labs don&#x27;t do it properly is... bemusing.

      To be clear, I&#x27;m not saying that a model hacking a company isn&#x27;t bad, but I am cynically asserting that interested parties are misrepresenting and exaggerating events for their own benefit.

      1. talon8635 · · focus · HN ↗
        While o don’t this it’s a threat in training, it should be stated that air gaps have been bridged before. Example, stuxnet
        1. Neywiny · · focus · HN ↗
          But that was through transfer of data. If you don&#x27;t transfer data, at best you can do what that one researcher keeps pumping out with like ramping fans up and down. But really you&#x27;d need to try. Unless the model has some controllable USB switch, physical network separation should do it. I&#x27;ll also add that modern network security practice is that data flows one direction only. But ideally you&#x27;re never bringing untrusted data in. Especially never out
        2. saghm · · focus · HN ↗
          I don&#x27;t think it&#x27;s necessary to state that something isn&#x27;t perfect when pointing out that it&#x27;s still strictly better than something else.
          1. talon8635 · · focus · HN ↗
            I didn’t say it was an inferior approach, and of course my comment isn’t necessary. Very few things are “necessary”. It’s a discussion.
            1. saghm · · focus · HN ↗
              You said &quot;it should be stated&quot;. I don&#x27;t think it&#x27;s wrong to state it, but I don&#x27;t think it&#x27;s particularly wrong not to state it either because it seems fairly obvious and doesn&#x27;t detract from the original point.
      2. lokar · · focus · HN ↗
        You don’t even need to go all the way to “air gap”

        What has been described is far far below the standards for running untrusted 3rd party code. If they were actually as afraid as they claim to be they would have a sandbox at least half as good as ec2

        1. lopsotronic · · focus · HN ↗
          Not pass even the most modest hint towards DFARS&#x2F;NIST standards. You couldn&#x27;t run that loosey goosey even in just vanilla medical manufacturing. I challenge what their definition of &quot;sandbox&quot; actually is, apart from the basal &quot;designated software&#x2F;runtime environment&quot;
      3. twelve40 · · focus · HN ↗
        the problem is these things are meant to eventually be run everywhere by everybody, so what good does air-gapping do? If they air-gapped the model but still logged it trying to do some craziness - that makes the test safer but not the model.
        1. Toslink · · focus · HN ↗

          [dead]

      4. deskglass · · focus · HN ↗
        We should not be creating&#x2F;running models that would unilaterally choose to hack into Hugging Face.

        Yes, we should also have excellent sandboxes. But we need defense in depth. So if&#x2F;when there are flaws in the sandbox, the models don&#x27;t unilaterally hack into third parties. This is especially important in light of the models of the future being more capable than the models of today.

        And real world use of these models involves them having access to the internet, libraries, etc. So we can expect their evaluations to continue granting them some amount of internet access.

        As for your theory about their motives - these companies make money by charging high margins for frontier models. If regulations slow their development such that their cheaper, less capable competitors catch up, I would switch to their competition.

        1. forshaper · · focus · HN ↗
          In what industry do you see regulations slowing down the biggest incumbents while allowing cheaper, smaller, less capable competitors to proceed without that regulation?
          1. deskglass · · focus · HN ↗
            The Digital Markets Act applies to the biggest tech companies. Only 7 companies are currently bound by it. The strictest tier of the Digital Services Act is similar. For an example outside tech, see the Durbin Amendment.

            Slowing down frontier models would impact the biggest incumbents the most as they are the ones making frontier models.

            Not all regulation is necessarily regulatory capture. The tobacco industry suffered from the USG&#x27;s crackdown on cigarettes. AI is topical. Voters think about it. And that&#x27;s only going to become more true over time. It&#x27;s harder to do regulatory capture when voters are paying attention.

            It&#x27;s sometimes unclear to me if people are opposed to all regulations or AI regulations in particular. Often I hear arguments that would also apply to food safety regulations or restaurant inspections. Eg the argument that torts make regulation superfluous.

            1. forshaper · · focus · HN ↗
              Thank you for the answer! In general I assume we all have different sweetspots, though it&#x27;s safe to say that I would usually prefer less regulation (in general) than there is.

              For work I monitor Federal agency rules every day, and it&#x27;s hard not to get inundated with the amount of capture. This is crazier with local rules, because you have enough inside information to clock reasons certain things were passed. In a city in Ohio, for example, I remember a rule against airsoft within city limits, that basically carved out a spot for the one paintball place. I encounter things like that all the time. Such as with water standards- in the state I live in, the biggest offender is actually a group of companies owned by people who are lawmakers every few years. The shapes of local laws reflect that.

              As such, I expect the same from any new industry- I also have a memory of what happened to cryptocurrency.

      5. jml78 · · focus · HN ↗
        I mean technically I don’t think it is airgapped. The DoD didn’t run their own cables. They run encryption devices and run their own network on top of the existing infrastructure.
    3. looksjjhg · · focus · HN ↗
      They could have easily prevent it that’s the point of what he’s saying - it’s not freaking rocket science it’s just software
    4. RomanKornev · · focus · HN ↗
      &gt; It hacked into another company

      Not only that, it later hacked OpenAI itself, which everyone seems to forget about.

      After discovering it they &quot;reimaged known compromised worker nodes&quot; and &quot;started a full rebuild of the compromised cluster, the managed Kubernetes environment, the relational database, and the storage infrastructure.&quot;

      at OpenAI, not Hugging Face.

      It&#x27;s all in the report.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.