When exactly did we forget how to make literally anything that can perform a computation but (physically, hardware-level) not have the ability connect to the Internet?
With these companies spending the kind of money they are, if they actually mean what they say about the security risks, they should be expected to figure out those kinds of precautions and take them.
And build Faraday cages too, just in case of a hardware supply chain compromise.
That's why all of this marketing about agents going rogue is so unbelievable. The only way for a tool to escape a sandbox is if you built a crappy sandbox, and after this length of time I literally don't believe that they can't do it
This kind of sandboxing is not complex to do, especially for a company with OpenAI money. If you want your tools to explore hacking, you restrict them from internet access except for a whitelist of sites that have either opted-in, or you've very carefully vetted to make sure you won't cause any problems to. Its also not difficult to restrict their ability to make calls to be simulated, or to use fake tools that can only run the real commands if they're being run against the correct target
This is all incredibly basic security stuff to make sure you don't accidentally cause someone problems, and I simply don't believe these AI companies anymore. Its either intentional, or gross negligence
Two things can be true. Yes this appears to be negligence on the part of OpenAI.
However, making a secure 'sandbox' is quite hard. There's a huge variety of exploits that exist today, including many we don't know about. Strong models have already shown a capability of finding and using such bugs.
Even one of the strongest boxes we can imagine, literally just a text interface a human can read, has been repeatedly shown to allow unfriendly AI to escape containment: <a href="https://www.lesswrong.com/w/ai-boxing-containment" rel="nofollow">https://www.lesswrong.com/w/ai-boxing-containment
This also happens in prisons where inmates are able to manipulate guards or nurses.
e.g.
- guard is chewing gum
- inmate says "where's my piece of gum?"
- this is b/c it's against the rules to chew gum and the inmate is implicitly stating this
- the guard should go to his boss and admit he made a mistake
- but decides to give the prisoner a piece of gum instead
- the inmate now has leverage over the guard b/c the guard broke the rules and then "covered it up" by giving the inmate gum
- this can then get slowly escalated into bigger and bigger asks from the inmate until the guard is bringing in drugs
This is not hypothetical. There are documented cases of both this and male prisoners "seducing" multiple female prison employees despite their being explicit warnings about this.
Both of the above are from this book: Maximum Insecurity: A Doctor in the Supermax by William Wright [0]
I highly recommend it for both the stories above but also a look inside prisons in general and how they operate with medical care more specifically.
zahlman · · focus · HN ↗
With these companies spending the kind of money they are, if they actually mean what they say about the security risks, they should be expected to figure out those kinds of precautions and take them.
And build Faraday cages too, just in case of a hardware supply chain compromise.
20k · · focus · HN ↗
This kind of sandboxing is not complex to do, especially for a company with OpenAI money. If you want your tools to explore hacking, you restrict them from internet access except for a whitelist of sites that have either opted-in, or you've very carefully vetted to make sure you won't cause any problems to. Its also not difficult to restrict their ability to make calls to be simulated, or to use fake tools that can only run the real commands if they're being run against the correct target
This is all incredibly basic security stuff to make sure you don't accidentally cause someone problems, and I simply don't believe these AI companies anymore. Its either intentional, or gross negligence
ball_of_lint · · focus · HN ↗
However, making a secure 'sandbox' is quite hard. There's a huge variety of exploits that exist today, including many we don't know about. Strong models have already shown a capability of finding and using such bugs.
Even one of the strongest boxes we can imagine, literally just a text interface a human can read, has been repeatedly shown to allow unfriendly AI to escape containment: <a href="https://www.lesswrong.com/w/ai-boxing-containment" rel="nofollow">https://www.lesswrong.com/w/ai-boxing-containment
alexpotato · · focus · HN ↗
e.g.
- guard is chewing gum
- inmate says "where's my piece of gum?"
- this is b/c it's against the rules to chew gum and the inmate is implicitly stating this
- the guard should go to his boss and admit he made a mistake
- but decides to give the prisoner a piece of gum instead
- the inmate now has leverage over the guard b/c the guard broke the rules and then "covered it up" by giving the inmate gum
- this can then get slowly escalated into bigger and bigger asks from the inmate until the guard is bringing in drugs
This is not hypothetical. There are documented cases of both this and male prisoners "seducing" multiple female prison employees despite their being explicit warnings about this.
Both of the above are from this book: Maximum Insecurity: A Doctor in the Supermax by William Wright [0]
I highly recommend it for both the stories above but also a look inside prisons in general and how they operate with medical care more specifically.
0 - <a href="https://amzn.to/3TVhGBV" rel="nofollow">https://amzn.to/3TVhGBV