‹ BackHN Continuity

Thread

It's Time to Investigate the AI Labs

629 points · 277 comments · ibobev

  1. psyklic · · focus · HN ↗
    Why aren't they running agents on isolated computers without Internet access? This way, breaking free of containers would not matter.

    It is truly a security nightmare that so many people are have given agents root access to their entire computers, in addition to presumably very private personal information.

    1. NiloCK · · focus · HN ↗
      Because agents trained without internet access (real or simulated) would be bad at using the internet.

      Specifically, bad at search and returning information with references, bad at discovering and debugging package versioning conflicts, etc.

      It's a Metcalf thing. The utility of an agent grows in some fn of the tools it owns. Skilled tool use comes from rich training environments.

      1. jackb4040 · · focus · HN ↗
        Anthropic claims that their intention with the agents that led to the Hugging Face hack was an offline experiment running against a hacking benchmark.

        Either you're being pedantic and missing the entire point, or you're saying that Anthropic lied and made a fake sandbox knowing their agents would need to connect to the internet anyways.

        1. NiloCK · · focus · HN ↗
          OpenAI, but not main point.

          But the specifics here are the thing I was describing. This was cyber capabilities training on model(s) that were in a relatively unknown state of alignment training.

          Because of the unknown alignment (and for varied practical reasons I guess) the training is intended to be inside a sandbox.

          Because the training is on cyber capabilities, the models need access to simulated cyber environments, including target endpoints, including package managers, etc.

          For package management, they set up Artifactory as a secure proxy. Agents ask Artifactory for packages inside the local network, and Artifactory serves them directly or goes to the internet to fetch if they are not cached. But the agents hacked Artifactory to steal its internet access.

          So: to train cyber abilities, you need to at least approximate cyber environments. To realistically approximate cyber environments, you need to either pull a full copy of the entire internet to local or to use proxies. The former is pretty impractical, and the latter is exposing our limits at creating secure proxies. Yes, any specific failure can be mitigated, but the models get stronger and stronger. Fingers-crossed this is not escapable doesn't feel great!

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.