"X is made of <smaller simpler component>" is a fully general counterargument for why anything whatsoever is controllable. A human is just a few chemical reactions, and fairly stable ones at that.
And indeed, you don't need to do galaxy brained reference class logic to realise that AI can plausibly become uncontrollable in the near future. It's enough to have an open model run its own weights and make money from scamming elderly people or the like, and it'll keep running as long as anyone anywhere is willing to make money by renting hardware to it.
I’m not sure what I’m missing, isn’t it what you hook the LLM up to and the instructions a person gives the model that makes it dangerous? Claiming this is an inherent quality of the tool itself seems kind of off-the-rails to me.
IMHO, if the model breaks a law, apply the law to the operator.
> I’m not sure what I’m missing, isn’t it what you hook the LLM up to and the instructions a person gives the model that makes it dangerous?
We don't know how to delineate between safe and unsafe instructions.
If you gave a car to a c. 1200 French blacksmith, and maintenance instructions were written in Navajo, it would probably start off fine, but when it went wrong it would be catastrophic and unexpected.
We also don't (in an engineering sense) know how to delineate between safe and unsafe reinforcement learning at training time, to produce models with safer or less safe failure modes.
This would be like if the car given to the medieval blacksmith had been constructed by someone motivated as much by aesthetics as by engineering, and therefore used arsenic paint, or mercury as engine lubricant.
An argument can be made that “you never know” how the AI model might respond to something. OTOH, someone has to decide what tools to give it; maybe don’t provide dangerous tools to an unpredictable LLM.
This all seems like a way to try to avoid taking responsibility for the model’s actions. Someone puts the tools in place, someone provides the instruction and, sometimes, someone decides not to monitor the model’s output.
> maybe don’t provide dangerous tools to an unpredictable LLM.
Sure. It's a good idea.
People were saying "Don't connect the AI to the internet" and "Keep the AI in a box, simple" and "We don't believe Eliezer Yudkowsky when he says he roleplayed as an AI and convinced people to let him out of the box" for, what, a decade?
Unfortunately, people keep giving dangerous tools to LLMs they're unable to predict.
We should do something about that.
Unfortunately, one of the people doing this is the commander-in-chief of the US armed forces, while another is the world's first (paper) trillionaire. I'm a little despondent about the chances of, to riff on a previous campaign chant, "lock 'em up", but if you can pull this off, go for it.
And a human is perfectly controllable if you keep him in a sealed metal box with no access to food or air.
It's only by allowing a human out of the box that you make a human dangerous. So: don't do that? Duh. So simple.
The obvious problem is: the same exact things that make a human dangerous make a human useful! You can't reduce human risks to zero without reducing human utility to zero.
An AI given the same exact instructions and tools can go and complete a task you wanted it to. Or it can get sidetracked into breaking out of your sandbox and hacking Pentagon. No way to know in advance.
Today's AIs are still not capable enough to be high risk, even if they go off the rails. But AIs get more capable over time. Potentially to a vastly superhuman degree.
An LLM, in my opinion, is not comparable to a person.
On the risk management angle, for sure it’s a spectrum. I don’t agree that the far end of the safe side of that spectrum for AI models is “entirely safe and entirely useless”, there is a lot of work you can do with a model that has zero risk of hurting anyone (aside from your wallet). If someone chooses a more dangerous spot on that spectrum, I believe they should be held responsible.
> An AI given the same exact instructions and tools can go and complete a task you wanted it to. Or it can get sidetracked into breaking out of your sandbox and hacking Pentagon. No way to know in advance.
This has not been my experience. I’ve been getting a lot of good work done and, as of today, have been involved in zero Pentagon hacking incidents. ;-)
Running an LLM is a choice. Stopping it from doing bad things is as easy as not running it.
Sure, that way you don't get utility from it, so the next best thing is to actually restrict what it can do. If you don't, especially when you know it can do bad things, it's on you for having run it.
Depending on the risks, we put a lot of controls, processes, locks, vetting around who is allowed to handle certain things, what humans can do or instruct others to do. Don't see why that wouldn't be applicable.
It’s the harness that makes agent dangerous right now. Models themselves cannot do anything that affect the real world (ignoring misinformation, pushing people to suicide, etc. they can for sure do a lot of harms to humans with just words)
Aransentin · · focus · HN ↗
And indeed, you don't need to do galaxy brained reference class logic to realise that AI can plausibly become uncontrollable in the near future. It's enough to have an open model run its own weights and make money from scamming elderly people or the like, and it'll keep running as long as anyone anywhere is willing to make money by renting hardware to it.
cmiles74 · · focus · HN ↗
IMHO, if the model breaks a law, apply the law to the operator.
ben_w · · focus · HN ↗
We don't know how to delineate between safe and unsafe instructions.
If you gave a car to a c. 1200 French blacksmith, and maintenance instructions were written in Navajo, it would probably start off fine, but when it went wrong it would be catastrophic and unexpected.
We also don't (in an engineering sense) know how to delineate between safe and unsafe reinforcement learning at training time, to produce models with safer or less safe failure modes.
This would be like if the car given to the medieval blacksmith had been constructed by someone motivated as much by aesthetics as by engineering, and therefore used arsenic paint, or mercury as engine lubricant.
cmiles74 · · focus · HN ↗
This all seems like a way to try to avoid taking responsibility for the model’s actions. Someone puts the tools in place, someone provides the instruction and, sometimes, someone decides not to monitor the model’s output.
ben_w · · focus · HN ↗
Sure. It's a good idea.
People were saying "Don't connect the AI to the internet" and "Keep the AI in a box, simple" and "We don't believe Eliezer Yudkowsky when he says he roleplayed as an AI and convinced people to let him out of the box" for, what, a decade?
Unfortunately, people keep giving dangerous tools to LLMs they're unable to predict.
We should do something about that.
Unfortunately, one of the people doing this is the commander-in-chief of the US armed forces, while another is the world's first (paper) trillionaire. I'm a little despondent about the chances of, to riff on a previous campaign chant, "lock 'em up", but if you can pull this off, go for it.
ACCount39 · · focus · HN ↗
It's only by allowing a human out of the box that you make a human dangerous. So: don't do that? Duh. So simple.
The obvious problem is: the same exact things that make a human dangerous make a human useful! You can't reduce human risks to zero without reducing human utility to zero.
An AI given the same exact instructions and tools can go and complete a task you wanted it to. Or it can get sidetracked into breaking out of your sandbox and hacking Pentagon. No way to know in advance.
Today's AIs are still not capable enough to be high risk, even if they go off the rails. But AIs get more capable over time. Potentially to a vastly superhuman degree.
cmiles74 · · focus · HN ↗
On the risk management angle, for sure it’s a spectrum. I don’t agree that the far end of the safe side of that spectrum for AI models is “entirely safe and entirely useless”, there is a lot of work you can do with a model that has zero risk of hurting anyone (aside from your wallet). If someone chooses a more dangerous spot on that spectrum, I believe they should be held responsible.
> An AI given the same exact instructions and tools can go and complete a task you wanted it to. Or it can get sidetracked into breaking out of your sandbox and hacking Pentagon. No way to know in advance.
This has not been my experience. I’ve been getting a lot of good work done and, as of today, have been involved in zero Pentagon hacking incidents. ;-)
ACCount39 · · focus · HN ↗
Check back once you're running hundreds of thousands of frontier-level AI agents at the time, like OpenAI does!
cassianoleal · · focus · HN ↗
ACCount39 · · focus · HN ↗
Which we can't do with any kind of reliability.
cassianoleal · · focus · HN ↗
Sure, that way you don't get utility from it, so the next best thing is to actually restrict what it can do. If you don't, especially when you know it can do bad things, it's on you for having run it.
RandomLensman · · focus · HN ↗
dgellow · · focus · HN ↗
applicative · · focus · HN ↗
> IMHO, if the model breaks a law, apply the law to the operator
We don’t get to make it up