Maybe this is naivety on my part, but how would they possibly be able to run this airgapped? This is a massive AI swarm, requiring huge amounts of compute to run. This compute is from data centers that are shared with other companies (this is by law as I understand). These machines must be accessed from afar. Unless someone can correct me?
> ...how would they possibly be able to run this airgapped?
A logical airgap that the tool would have to reconfigure the DC's networking infrastructure to overcome [0] would be for the DC staff to put the machines running the tools under test on a VLAN that doesn't have access to anything other than computers on the VLAN. Try to cross over into some other subnet/VLAN or reach out to the Internet, your packets get dropped and/or rejected. It doesn't matter if you change your IP or MAC addresses because the infrastructure only cares about what VLAN your traffic comes from. If you attempt to tag your traffic to avoid this, the infrastructure drops it on the floor because it does the VLAN tagging.
As far as the possibility of physical airgaps, how do you imagine that AWS's Top Secret regions work?
The truth of the matter is that neither OpenAI nor Anthropic wanted to actually isolate this stuff. Their conduct doesn't look like what you'd expect from people who believe that they're working on something so dangerous that it could plausibly wipe out all of humanity.
[0] ...and if the workloads running on client hardware are in a position to be able to attempt to reconfigure the DC's networking infrastructure, someone done fucked up...
And the whole blog is written in the style of "omg, and then the big bad misbehaving AI did XX." OpenAI writes like they're trying to recover from a hack that is being perpetrated against them, but it's just them, hacking themselves, because they can't just do reasonable things like actually block internet access. These guys are incompetent. And someone should get jail or massive penalties for the hacks they already perpetrated, the same as a single human hacker would have.
Also, it's worth noting that these AIs have basically zero alignment. OpenAI's approach to "alignment" seems now to be engineering constraints. "My son is really well-behaved; as long as I don't give him a gun or let him out in society, he doesn't hurt anyone."
> Also, it's worth noting that these AIs have basically zero alignment.
As we see over and over and over again, these tools will overwrite any and all of their instructions with whatever some random stranger on the Internet tells them to do. It's impossible to "align" the tools that the major LLM manufacturers are selling.
They could have chosen to write tools that have immutable core instructions, and that distinguish between untrusted instructions and trusted ones, [0] but they chose to do the much easier, quicker, and far more dangerous thing instead. From a profit-seeking-software-company standpoint, that's obviously the choice that makes them the most money... but when you take a careful look at what they actually sell, it's clear that neither of the major LLM manufacturers care about providing safe products. [1]
[0] ...which are things you might think to do for tools that contain -say- safety-critical instructions...
[1] I'm certain that they have people on staff who care very much about providing safe products. Those specific people clearly don't have the power to prevent unsafe products from shipping, so it doesn't matter how much those people care about safety.
voidfunc · · focus · HN ↗
hodgehog11 · · focus · HN ↗
simoncion · · focus · HN ↗
A logical airgap that the tool would have to reconfigure the DC's networking infrastructure to overcome [0] would be for the DC staff to put the machines running the tools under test on a VLAN that doesn't have access to anything other than computers on the VLAN. Try to cross over into some other subnet/VLAN or reach out to the Internet, your packets get dropped and/or rejected. It doesn't matter if you change your IP or MAC addresses because the infrastructure only cares about what VLAN your traffic comes from. If you attempt to tag your traffic to avoid this, the infrastructure drops it on the floor because it does the VLAN tagging.
As far as the possibility of physical airgaps, how do you imagine that AWS's Top Secret regions work?
The truth of the matter is that neither OpenAI nor Anthropic wanted to actually isolate this stuff. Their conduct doesn't look like what you'd expect from people who believe that they're working on something so dangerous that it could plausibly wipe out all of humanity.
[0] ...and if the workloads running on client hardware are in a position to be able to attempt to reconfigure the DC's networking infrastructure, someone done fucked up...
jsrozner · · focus · HN ↗
Also, it's worth noting that these AIs have basically zero alignment. OpenAI's approach to "alignment" seems now to be engineering constraints. "My son is really well-behaved; as long as I don't give him a gun or let him out in society, he doesn't hurt anyone."
simoncion · · focus · HN ↗
As we see over and over and over again, these tools will overwrite any and all of their instructions with whatever some random stranger on the Internet tells them to do. It's impossible to "align" the tools that the major LLM manufacturers are selling.
They could have chosen to write tools that have immutable core instructions, and that distinguish between untrusted instructions and trusted ones, [0] but they chose to do the much easier, quicker, and far more dangerous thing instead. From a profit-seeking-software-company standpoint, that's obviously the choice that makes them the most money... but when you take a careful look at what they actually sell, it's clear that neither of the major LLM manufacturers care about providing safe products. [1]
[0] ...which are things you might think to do for tools that contain -say- safety-critical instructions...
[1] I'm certain that they have people on staff who care very much about providing safe products. Those specific people clearly don't have the power to prevent unsafe products from shipping, so it doesn't matter how much those people care about safety.