When exactly did we forget how to make literally anything that can perform a computation but (physically, hardware-level) not have the ability connect to the Internet?
With these companies spending the kind of money they are, if they actually mean what they say about the security risks, they should be expected to figure out those kinds of precautions and take them.
And build Faraday cages too, just in case of a hardware supply chain compromise.
So you build an offline tool that simulates it, or you proxy through your own service where you can ratelimit, inspect, and restrict the traffic
None of this is difficult to do, and its impossible to believe that a company the scale of OpenAI doesn't know this. I've built web crawlers and scrapers before, and the thing you do is test them extensively offline against simulated versions of the sites in question, and then very VERY cautiously run them against the prod versions so that you don't cause anyone any issues
The only reason not to do this is because OpenAI doesn't give a rats ass about the internet as a public good, nor the legal consequences of compromising systems
If you own a network and the servers, you can DPI every single packet and see literally every bit of information. All of the text to the "forum" that they created must have been in -outbound- packets to their compromised package manager, by definition. If they can't properly analyze network traffic, they should not be running 'sandboxes'.
Anything beyond baseline would be observable- silence, malformed packets, too much egress, unusually large packets, etc
zahlman · · focus · HN ↗
With these companies spending the kind of money they are, if they actually mean what they say about the security risks, they should be expected to figure out those kinds of precautions and take them.
And build Faraday cages too, just in case of a hardware supply chain compromise.
serbuvlad · · focus · HN ↗
20k · · focus · HN ↗
None of this is difficult to do, and its impossible to believe that a company the scale of OpenAI doesn't know this. I've built web crawlers and scrapers before, and the thing you do is test them extensively offline against simulated versions of the sites in question, and then very VERY cautiously run them against the prod versions so that you don't cause anyone any issues
The only reason not to do this is because OpenAI doesn't give a rats ass about the internet as a public good, nor the legal consequences of compromising systems
serbuvlad · · focus · HN ↗
It literally seems like they are doing just that, and the agents are just finding holes in that.
ceroxylon · · focus · HN ↗
Anything beyond baseline would be observable- silence, malformed packets, too much egress, unusually large packets, etc