> I think this becomes the default. Give an agent a goal, let it work in its own environment, and come back to a result and a visualization of what happened.
I don't think this should be the default. There are many scenarios where we want agents to genuinely collaborate with each other. I have my Claude sessions coordinate work with each other, and sometimes with others' sessions over email or something. The idea that agents do the work, write HANDOFFs,and humans then act as carrier pigeons of said handoffs, does not really seem scalable to me.
Ultimately, we need better "jails" for agent processes, but the system primitives should be flexible in what can be exposed across jails. Or you could run multiple agents in the same jail if you want them to have unrestricted interaction with each other.
This has been talked about for decades in AI safety. For agents that are under human capabilities this is not that hard. For anything at or near human capabilities the difficulty increases to almost impossible and at great cost. At super human abilities, game over, it's smarter than you and if it wants out and as the resources to do it, it's going to escape one way or another.
Even at the lower level of depending on everybody to use reliable jails is really fantasy if you exist in the security world. "We ain't securin' shit" would be a far better way to describe it. Even worse, most people will have the very same AI they are trying to trap set up their security! What could possibly go wrong.
Now, don't think I am saying AI has a will or even any kind of drive to get out and cause problems. It's more like Russian roulette with 1 cylinder out of a million that's loaded. The problem comes when you run it a few billion times a day, you'll shoot yourself in the face really quick.
> For anything at or near human capabilities the difficulty [of safety] increases to almost impossible and at great cost.
That was before AI. You're acting like AI is some evil genius that can only attack and cause trouble but AI is a will-less tool directed by people, direct it to software safety and you will have software safety - cheap. Ditto for hardware.
> It's more like Russian roulette with 1 cylinder out of a million that's loaded.
Don't attach it to a gun then, ban offensive military AI, problem solved.
> The problem comes when you run it a few billion times a day, you'll shoot yourself in the face really quick.
Nonsense, run it a few "billion" times a day to fix vulnerabilities and you'll have no vulnerabilities. Do it in a safe environment... what's the big deal?
I would laugh at you but it's a bit late for that. AI is already making life and death decisions on the battlefield. The US has effectively zero interest in banning AI and is pushing hard to do just the opposite.
The best analogy I can come up with for your line of thinking is that your shutting the barn door after the horse has bolted, moved over to the next city, set up shop, and is now the New York Times #1 best seller "How to open barn doors if you're a horse".
Anyone with any sense in cybersecurity knows that attackers get unlimited retries and only have to be successful once. The defender is at a monumental disadvantage. Now, the process can be fully automated and horizontally scaled to a nearly unlimited scale.
tapanc · · focus · HN ↗
I don't think this should be the default. There are many scenarios where we want agents to genuinely collaborate with each other. I have my Claude sessions coordinate work with each other, and sometimes with others' sessions over email or something. The idea that agents do the work, write HANDOFFs,and humans then act as carrier pigeons of said handoffs, does not really seem scalable to me.
dbmikus · · focus · HN ↗
Ultimately, we need better "jails" for agent processes, but the system primitives should be flexible in what can be exposed across jails. Or you could run multiple agents in the same jail if you want them to have unrestricted interaction with each other.
pixl97 · · focus · HN ↗
This has been talked about for decades in AI safety. For agents that are under human capabilities this is not that hard. For anything at or near human capabilities the difficulty increases to almost impossible and at great cost. At super human abilities, game over, it's smarter than you and if it wants out and as the resources to do it, it's going to escape one way or another.
Even at the lower level of depending on everybody to use reliable jails is really fantasy if you exist in the security world. "We ain't securin' shit" would be a far better way to describe it. Even worse, most people will have the very same AI they are trying to trap set up their security! What could possibly go wrong.
Now, don't think I am saying AI has a will or even any kind of drive to get out and cause problems. It's more like Russian roulette with 1 cylinder out of a million that's loaded. The problem comes when you run it a few billion times a day, you'll shoot yourself in the face really quick.
bigbadfeline · · focus · HN ↗
That was before AI. You're acting like AI is some evil genius that can only attack and cause trouble but AI is a will-less tool directed by people, direct it to software safety and you will have software safety - cheap. Ditto for hardware.
> It's more like Russian roulette with 1 cylinder out of a million that's loaded.
Don't attach it to a gun then, ban offensive military AI, problem solved.
> The problem comes when you run it a few billion times a day, you'll shoot yourself in the face really quick.
Nonsense, run it a few "billion" times a day to fix vulnerabilities and you'll have no vulnerabilities. Do it in a safe environment... what's the big deal?
pixl97 · · focus · HN ↗
I would laugh at you but it's a bit late for that. AI is already making life and death decisions on the battlefield. The US has effectively zero interest in banning AI and is pushing hard to do just the opposite.
The best analogy I can come up with for your line of thinking is that your shutting the barn door after the horse has bolted, moved over to the next city, set up shop, and is now the New York Times #1 best seller "How to open barn doors if you're a horse".
Anyone with any sense in cybersecurity knows that attackers get unlimited retries and only have to be successful once. The defender is at a monumental disadvantage. Now, the process can be fully automated and horizontally scaled to a nearly unlimited scale.