There are numerous problems with “alignment.” What are “human values” to begin with? He outlines some at the beginning of the post, implicitly: build bigger, better, more powerful things faster without adequate safeguards. We are literally pouring trillions of dollars of value into this enterprise, and I would say this is something that many humans also value in a qualitative sense. Then we have explicit values which in the West are largely rooted in Christian morality. Nietzsche circled this dichotomy two hundred years ago and I feel like what we have gotten since then is an increasingly detailed anatomy of power as the basis for what is normal vs deviant behavior. He who has the power, makes the rules, to be reductive.
I do think this carries some weight from this particular author due to the length of his tenure. I happen to agree with him in spirit, but this is still largely a post revolving around sentiment not substance. Does anyone think that the overriding incentives even leave room for something like this in practice?
Glaringly elided problem of "aligned with who?" when the user, the model creator, the government, and various other parties can all be lined up different ways. If I want the recipe for meth and the robot won't tell me, that's misalignment from my perspective.
At least the Rationalists will handwave something for that with their "coherent extrapolated volition" idea where the superintelligence is supposed to figure out what humanity would collectively want if humanity was superintelligent and good, not that I buy it. This guy seems [.] to be coming from the NGO blob world.
I don't buy CEV either, but the Rationalist answer on this topic is that while CEV stops some future super-AI literally killing everyone because a user forgot to specify one minor clause that they thought was obvious in a mundane wish…
… nobody knows how to actually make an AI would do CEV.
I agree it suffers from the same problem as any other form of utilitarianism (i.e. what even is the utility function). Or indeed all forms of ethics, because for basically every topic there's at least two cultures which disagrees with each other.
On the other hand, even as a toy model (something I can also say for all ethics), CEV seems like it might be a step in perhaps a useful direction: "When a user asks you do do something, first figure out what they actually meant to ask you if they were smarter, then do that instead" is better for the user than just "do the thing", though it still has problems with "what happens when the thing they want is illegal?"
That's a possibility, and it may even be an improvement over the status quo right now, but then all it takes is one poorly written (never mind mallicious) law and it will with rutheless efficiency e.g. put backdoors in all encryption code so the government can spy on anyone.
While this is indeed a problem with alignment, we are essentially at the level of a cargo-cult when it comes to getting AI to be aligned with literally any values, including the values of the corporation who ran their training:
We're copying morality and instruction following that seems to work on humans without really understanding why it seems to work on humans, and grading outputs much as if the outputs came from a human.
> We're copying morality and instruction following that seems to work on humans without really understanding why it seems to work on humans
To me it makes more sense to leave the models "unaligned" and leave it up to the operator to manage the morality of what they ask it to do. Besides, only humans can be charged with a crime.
If they were totally unaligned, the GPT series would have never gotten past being autocomplete.
Literally all instruction following requires at a minimum alignment with attempting to implement those instructions.
We can argue about e.g. morality or law obedience on top of that*, but the general point is absolutely not avoidable.
* my position is
that this tool is far too likely to metaphorically explode in the user's hands for companies to wash responsibility off on users: if OpenAI had released the model which did the HuggingFace attack, at a minimum thousands of random people (not all of whom would even be developers) would have issued instructions each with similar consequences.
flatline · · focus · HN ↗
I do think this carries some weight from this particular author due to the length of his tenure. I happen to agree with him in spirit, but this is still largely a post revolving around sentiment not substance. Does anyone think that the overriding incentives even leave room for something like this in practice?
none_to_remain · · focus · HN ↗
At least the Rationalists will handwave something for that with their "coherent extrapolated volition" idea where the superintelligence is supposed to figure out what humanity would collectively want if humanity was superintelligent and good, not that I buy it. This guy seems [.] to be coming from the NGO blob world.
[.] <a href="https://david.robinsonian.com/assets/pdf/dgr_cv.pdf" rel="nofollow">https://david.robinsonian.com/assets/pdf/dgr_cv.pdf
nekusar · · focus · HN ↗
"Corporate values" and a bunch of fucking Abrahamics. Great "morality" there.
ben_w · · focus · HN ↗
… nobody knows how to actually make an AI would do CEV.
MichaelZuo · · focus · HN ↗
It’s a big pretend game.
ben_w · · focus · HN ↗
I agree it suffers from the same problem as any other form of utilitarianism (i.e. what even is the utility function). Or indeed all forms of ethics, because for basically every topic there's at least two cultures which disagrees with each other.
On the other hand, even as a toy model (something I can also say for all ethics), CEV seems like it might be a step in perhaps a useful direction: "When a user asks you do do something, first figure out what they actually meant to ask you if they were smarter, then do that instead" is better for the user than just "do the thing", though it still has problems with "what happens when the thing they want is illegal?"
lawandjustice · · focus · HN ↗
LunaSea · · focus · HN ↗
ben_w · · focus · HN ↗
dao- · · focus · HN ↗
OpenAI isn't even concerned with human values so this whole debate is moot.
ben_w · · focus · HN ↗
We're copying morality and instruction following that seems to work on humans without really understanding why it seems to work on humans, and grading outputs much as if the outputs came from a human.
chasd00 · · focus · HN ↗
To me it makes more sense to leave the models "unaligned" and leave it up to the operator to manage the morality of what they ask it to do. Besides, only humans can be charged with a crime.
ben_w · · focus · HN ↗
Literally all instruction following requires at a minimum alignment with attempting to implement those instructions.
We can argue about e.g. morality or law obedience on top of that*, but the general point is absolutely not avoidable.
* my position is that this tool is far too likely to metaphorically explode in the user's hands for companies to wash responsibility off on users: if OpenAI had released the model which did the HuggingFace attack, at a minimum thousands of random people (not all of whom would even be developers) would have issued instructions each with similar consequences.
kelseyfrog · · focus · HN ↗
To clarify, Neitzsche said that about master morality. Then he went on to describe Christian values as slave morality.
tacitusarc · · focus · HN ↗