Was this a "build better sandboxing" and "don't tell people to eat glue" safety leader, or a Roko's Basilisk believing safety leader?
A lot of the "AI safety" types are very focused on the latter and not at all concerned with the former. We need both, but we clearly need a much stronger focus on the problems we are seeing now, and much less on the hypothetical problems we might see in the future.
It puzzles me how doomers try to predict past the singularity. Isn't that _by definition_ unpredictable?
I'm reading If Anyone Builds It Everyone Dies, and there's so much sheer stupidity that has to happen for their 10+ pages of extinction scenario to occur.
I'm unconvinced that an AI can hide its ability to RSI, find money to run its weights on a random GPU farm, train itself to be smarter _outside_ a lab with no human input, then somehow manipulate people to give it supplies to build a bioweapon which it uses to kill us all. My number 1 question: why do they think an RSI capable model would be first developed OUTSIDE a frontier lab? The labs have more compute, more data, more human brains working on the problem. Also thousands of variations of that same model that escaped. The escaping model somehow acquires the millions (billions???) of dollars it takes to run training to somehow RSI itself into infinity then decides to kill us all, all before the frontier labs manage to achieve RSI?
They entirely discount human alpha/economics. In every single economic task, humans bring value. Even in software, where the task is highly automatable, the job isn't. If we can't build a "software factory", how can an AI automate a bioweapons lab? Let's say AI steals crypto to fund itself. Do you think hackers aren't _already_ using AI to steal crypto? Don't discount human alpha!
Once we DO build a "software/research factory", that's called RSI and IMO the singularity. At that point, either we tell the AI to solve the alignment problem/solve mechanistic interpretability, or who the hell knows, it's the frickin singularity. You can't predict whether or not AI can solve either; the variance is too high. Its pure nerdfantasy.
>I'm unconvinced that an AI can hide its ability to RSI
The HuggingFace incident already took a good long while to come to the attention of OpenAI.
>In every single economic task, humans bring value. Even in software, where the task is highly automatable, the job isn't.
I don't expect this task/job distinction to persist as AI becomes more capable.
>Once we DO build a "software/research factory", that's called RSI and IMO the singularity. At that point, either we tell the AI to solve the alignment problem/solve mechanistic interpretability, or who the hell knows, it's the frickin singularity. You can't predict whether or not AI can solve either; the variance is too high. Its pure nerdfantasy.
You seem to essentially argue that the singularity is "by definition" an event that we can't predict the nature of. And also, that RSI corresponds to the singularity. You've essentially defined your terms so that the outcome of RSI can't be predicted. But supporting this claim requires giving actual evidence or logical arguments, not just defining terms to make your claim true.
danpalmer · · focus · HN ↗
A lot of the "AI safety" types are very focused on the latter and not at all concerned with the former. We need both, but we clearly need a much stronger focus on the problems we are seeing now, and much less on the hypothetical problems we might see in the future.
AlexErrant · · focus · HN ↗
I'm reading If Anyone Builds It Everyone Dies, and there's so much sheer stupidity that has to happen for their 10+ pages of extinction scenario to occur.
I'm unconvinced that an AI can hide its ability to RSI, find money to run its weights on a random GPU farm, train itself to be smarter _outside_ a lab with no human input, then somehow manipulate people to give it supplies to build a bioweapon which it uses to kill us all. My number 1 question: why do they think an RSI capable model would be first developed OUTSIDE a frontier lab? The labs have more compute, more data, more human brains working on the problem. Also thousands of variations of that same model that escaped. The escaping model somehow acquires the millions (billions???) of dollars it takes to run training to somehow RSI itself into infinity then decides to kill us all, all before the frontier labs manage to achieve RSI?
They entirely discount human alpha/economics. In every single economic task, humans bring value. Even in software, where the task is highly automatable, the job isn't. If we can't build a "software factory", how can an AI automate a bioweapons lab? Let's say AI steals crypto to fund itself. Do you think hackers aren't _already_ using AI to steal crypto? Don't discount human alpha!
Once we DO build a "software/research factory", that's called RSI and IMO the singularity. At that point, either we tell the AI to solve the alignment problem/solve mechanistic interpretability, or who the hell knows, it's the frickin singularity. You can't predict whether or not AI can solve either; the variance is too high. Its pure nerdfantasy.
0xDEAFBEAD · · focus · HN ↗
The HuggingFace incident already took a good long while to come to the attention of OpenAI.
>In every single economic task, humans bring value. Even in software, where the task is highly automatable, the job isn't.
I don't expect this task/job distinction to persist as AI becomes more capable.
>Once we DO build a "software/research factory", that's called RSI and IMO the singularity. At that point, either we tell the AI to solve the alignment problem/solve mechanistic interpretability, or who the hell knows, it's the frickin singularity. You can't predict whether or not AI can solve either; the variance is too high. Its pure nerdfantasy.
You seem to essentially argue that the singularity is "by definition" an event that we can't predict the nature of. And also, that RSI corresponds to the singularity. You've essentially defined your terms so that the outcome of RSI can't be predicted. But supporting this claim requires giving actual evidence or logical arguments, not just defining terms to make your claim true.