Why aren't they running agents on isolated computers without Internet access? This way, breaking free of containers would not matter.
It is truly a security nightmare that so many people are have given agents root access to their entire computers, in addition to presumably very private personal information.
How about they create a virus that copies itself onto USB drives? Getting that virus out there doesn't sound that hard, just creare some crap game.exe and put it on itch.io for free. All it takes is for a kid to borrow nuclear-engineer-dad's USB drive and plug it in their infected PC.
I'm not a fan of using AI, it doesn't amplify the work I do that much, but I'm terrified by what is already possible today.
> All it takes is for a kid to borrow nuclear-engineer-dad's USB drive
Do nuclear-engineer-dads take work usb drive home. Are their computer even allowed to have usb. I'm typing this on a dell latitude, and you can disable usb and other peripherals inside the bios.
My last job, I had to apply for an exemption to use the USB ports. And it was run of the mill software engineering.
...why not pitch a start up that sells AI agent that don't train with/for internet access/usage?
1) Because effectively all of the hardware used for that task is earmarked for the major LLM manufacturers.
2) Because in a just world, the major LLM manufacturers would get punished for their flagrant deception, lawbreaking, negligence, and recklessness, [0] go out of business, you'd get all the hardware that they were using at fire sale prices and save a ton of money compared to attempting to start up now.
[0] Based on their recent claims, building WMDs [1] without adequate safeguards is the most negligent and reckless thing they've been doing.
[1] It's fair to call anything with a 10% chance of wiping out all of humanity a WMD.
I feel like you missed my point. To be slightly more explicit, I find it comical how willing security people are to build a "secure" product that's completely useless for getting work done.
Your second clause is sufficient there: they're not going to have a good time trying to sell AI that can't access the internet, because accessing the internet (and other computer systems) makes AI much more useful.
Sooner or later, agents must be connected to the real world to do economically useful work. The sandbox failures are bad, sure, but the underlying misbehavior is the root problem here.
I think the way they're connected to the real world doesn't have to be as free-form as we're doing now.
If you think AI can help design new drugs, great! Let's pick some restricted DSL for describing molecules, preparation instructions, hook it up to the right set of lab robots that can pipette and centrifuge and whatever, _restrict its output to that DSL_, and let it iterate. The 'let an AI design drugs' process doesn't require that the agent can also, say, hack a niche German wiki.
For every area where we think these AIs can be _so valuable_ because they'll find solutions that humans wouldn't (or find them faster), then doesn't that value also imply it would be worth it for some humans to do some advanced setup of their specific domain-specific tools so it can be run in a sandbox equipped with only what it needs and not more?
Anthropic claims that their intention with the agents that led to the Hugging Face hack was an offline experiment running against a hacking benchmark.
Either you're being pedantic and missing the entire point, or you're saying that Anthropic lied and made a fake sandbox knowing their agents would need to connect to the internet anyways.
OpenAI was involved in the hugging face incident, not anthropic, and yes, inadequate sandboxing was a major factor, even the wikipedia page states this: <a href="https://en.wikipedia.org/wiki/OpenAI%E2%80%93HuggingFace_incident" rel="nofollow">https://en.wikipedia.org/wiki/OpenAI%E2%80%93HuggingFace_inc...
But the specifics here are the thing I was describing. This was cyber capabilities training on model(s) that were in a relatively unknown state of alignment training.
Because of the unknown alignment (and for varied practical reasons I guess) the training is intended to be inside a sandbox.
Because the training is on cyber capabilities, the models need access to simulated cyber environments, including target endpoints, including package managers, etc.
For package management, they set up Artifactory as a secure proxy. Agents ask Artifactory for packages inside the local network, and Artifactory serves them directly or goes to the internet to fetch if they are not cached. But the agents hacked Artifactory to steal its internet access.
So: to train cyber abilities, you need to at least approximate cyber environments. To realistically approximate cyber environments, you need to either pull a full copy of the entire internet to local or to use proxies. The former is pretty impractical, and the latter is exposing our limits at creating secure proxies. Yes, any specific failure can be mitigated, but the models get stronger and stronger. Fingers-crossed this is not escapable doesn't feel great!
I think there’s a 2022 way of thinking how models are trained/evaluated, and there’s 2026 way of doing things. Models that are not evaluated with public internet access are pretty useless nowadays for daily operations.
Because then it wouldn't scare the governments into regulating models that are undercutting them. This is just an obvious tactic too and it seems like it might be backfiring a bit. Congress might be a bit dumb, but they're not that dumb.
psyklic · · focus · HN ↗
It is truly a security nightmare that so many people are have given agents root access to their entire computers, in addition to presumably very private personal information.
malisper · · focus · HN ↗
They probably will soon, but even air-gapping may not be enough. See stuxnet for instance
DamnInteresting · · focus · HN ↗
I'll be impressed if AI agents find a way to get infected USB drives out into meatspace.
tikotus · · focus · HN ↗
I'm not a fan of using AI, it doesn't amplify the work I do that much, but I'm terrified by what is already possible today.
skydhash · · focus · HN ↗
Do nuclear-engineer-dads take work usb drive home. Are their computer even allowed to have usb. I'm typing this on a dell latitude, and you can disable usb and other peripherals inside the bios.
My last job, I had to apply for an exemption to use the USB ports. And it was run of the mill software engineering.
tikotus · · focus · HN ↗
olejorgenb · · focus · HN ↗
jppittma · · focus · HN ↗
simoncion · · focus · HN ↗
1) Because effectively all of the hardware used for that task is earmarked for the major LLM manufacturers.
2) Because in a just world, the major LLM manufacturers would get punished for their flagrant deception, lawbreaking, negligence, and recklessness, [0] go out of business, you'd get all the hardware that they were using at fire sale prices and save a ton of money compared to attempting to start up now.
[0] Based on their recent claims, building WMDs [1] without adequate safeguards is the most negligent and reckless thing they've been doing.
[1] It's fair to call anything with a 10% chance of wiping out all of humanity a WMD.
jppittma · · focus · HN ↗
micromacrofoot · · focus · HN ↗
aesthesia · · focus · HN ↗
specked-citrus · · focus · HN ↗
abeppu · · focus · HN ↗
If you think AI can help design new drugs, great! Let's pick some restricted DSL for describing molecules, preparation instructions, hook it up to the right set of lab robots that can pipette and centrifuge and whatever, _restrict its output to that DSL_, and let it iterate. The 'let an AI design drugs' process doesn't require that the agent can also, say, hack a niche German wiki.
For every area where we think these AIs can be _so valuable_ because they'll find solutions that humans wouldn't (or find them faster), then doesn't that value also imply it would be worth it for some humans to do some advanced setup of their specific domain-specific tools so it can be run in a sandbox equipped with only what it needs and not more?
srdjanr · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
NiloCK · · focus · HN ↗
Specifically, bad at search and returning information with references, bad at discovering and debugging package versioning conflicts, etc.
It's a Metcalf thing. The utility of an agent grows in some fn of the tools it owns. Skilled tool use comes from rich training environments.
jackb4040 · · focus · HN ↗
Either you're being pedantic and missing the entire point, or you're saying that Anthropic lied and made a fake sandbox knowing their agents would need to connect to the internet anyways.
voidhorse · · focus · HN ↗
NiloCK · · focus · HN ↗
But the specifics here are the thing I was describing. This was cyber capabilities training on model(s) that were in a relatively unknown state of alignment training.
Because of the unknown alignment (and for varied practical reasons I guess) the training is intended to be inside a sandbox.
Because the training is on cyber capabilities, the models need access to simulated cyber environments, including target endpoints, including package managers, etc.
For package management, they set up Artifactory as a secure proxy. Agents ask Artifactory for packages inside the local network, and Artifactory serves them directly or goes to the internet to fetch if they are not cached. But the agents hacked Artifactory to steal its internet access.
So: to train cyber abilities, you need to at least approximate cyber environments. To realistically approximate cyber environments, you need to either pull a full copy of the entire internet to local or to use proxies. The former is pretty impractical, and the latter is exposing our limits at creating secure proxies. Yes, any specific failure can be mitigated, but the models get stronger and stronger. Fingers-crossed this is not escapable doesn't feel great!
tokioyoyo · · focus · HN ↗
dawnerd · · focus · HN ↗