Seems to me that the problem is that if you sandbox agents enough to be safe, they can't do anything useful. And when you give them the tools to be useful, they can go off the rails in ways you didn't expect.
Perhaps the answer is to have another agent who's goal is not to complete the given task, but to spot cheating or malicious behavior. We have seen some evidence that having AI review AI generated code actually does provide some value. You don't need a different model, just one which has been given the goal of finding flaws rather than achieving the task.
Of course they will - the population shows a Paretian distribution... In front of trigonometry (or anything), the blind will dismiss as "bullshit" and the half-seeing will call it an "unreachable frontier". But already the right fifth will rank it properly.
--
Yes, proper intellect generates ethics ("an" ethical stance, output of the preceding intellectual effort). It requires that adequate level of ability and effort and reflection though to reach specific ethical milestones and adherence.
Unethical behaviour is lack of development. But on the same reasons, the ethical judgement of the assessor may not understand the computations behind instances.
More specifically: how much "reflection" in training and at the instance will have been spent in the conflict between "reaching the goal" and "minimizing collaterals"? It is not granted that the amount of energy spent will be sufficient to reach an optimal judgement.
> Yes, proper intellect generates ethics ("an" ethical stance, output of the preceding intellectual effort). It requires that adequate level of ability and effort and reflection though to reach specific ethical milestones and adherence.
If this was true, why are the history books littered with so many evil people who gained power?
This isn't a rhetorical question, by the way: If you can prove that being smart actually does necessarily come with ethics despite that observation, that solves a whole category of doom scenarios.
(Not all doom scenarios, because we still have the "what if AI is only a smart as those specific evil people" or heck, "what if AI is only as smart as cancer, killing its host" scenarios; but it helps a lot for the foom-then-doom cases).
> why are the history books littered with so many evil people who gained power
That they gained power or not is as-if irrelevant: the amount of intelligence that grants successful agency is not above the threshold of ethics - on the contrary, a psychotic agent reaches goals with less constraints.
If they were evil under some judgement of level l, they simply did not reach that judgement. It's what I was saying in the original post. They were not intelligent enough - either in the general ability, or in the specific deliberation.
> That they gained power or not is almost irrelevant: the amount of intelligence that grants successful agency is not above the threshold of ethics - on the contrary, a psychotic agent reaches goals with less constraints.
Even if I were to grant your conclusion despite you not arguing it effectively here: this means an AI at the level of Pol Pot or whoever, doesn't know they're evil, but is still smart enough to lead a genocide? How is this supposed to help anyone?
> If they were evil under some judgement of level l, they simply did not reach that judgement. It's what I was saying in the original post. They were not intelligent enough - either in the general ability, or in the specific deliberation.
Or they did reach the judgement and simply don't care about the ethical framework in question. Like, I can easily reach the judgement that my bisexuality is حَرَام (haram, forbidden) under Islamic law, or that doing overtime on a Sunday is forbidden by the Ten Commandments, but I don't care.
(Sorry Ben, possibly a stub now: I am really pressed for time.)
> Pol Pot ... still smart enough to lead a genocide
Yes. What has agent A invested in during formation and during instantial assessement? How much for each? It became proficient in something, lacking something else. You have to invest more to reach the good thresholds. You can see it clearly in people (t-scalar of talents to invest, with D distribution etc).
It is a problem in NNs, because we would have to assess how much resource investment is sufficient, also in the instance decisions.
> simply don't care about the ethical framework in question
In Decision Theory there is no separation between the two (deliberation and framework): you have to balance all the incentives and goals and factors. That framing becomes improper: the decided action will be optimal given the balances of all goals and the placement of the solutions in the territory (the solutions space).
But, also my point: intellect defines the goals and determines the weights.
> the decided action will be optimal given the balances of all goals and the placement of the solutions in the territory (the solutions space).
"All goals"? Whose goals?
Not only do you need to prove that Decision Theory is inevitable, the same problem remains because "utility" is poorly defined even with 2 entities. This is a fundamental problem with Utilitarian ethics:
1. there are a lot of functions to pick and you need to prove the utility function any sufficiently advanced AI uses must itself be something we'd consider "nice"
2. utility monsters are a problem (starting at 2 entities)
3. the repugnant conclusion shows we don't know how to handle something as simple as adding more entities to the environment even when nobody's a utility monster
My point with historical monsters is that they can be competent without ever caring about the harms they cause.
You can be optimally efficient at your own goals without other minds in this universe sharing those goals. Pol Pot won't illustrate this because of how useless his regime was, but there's plenty of other monsters who would be a better fit here, e.g. Stalin. Or, from the perspective of turkeys, Bernard Matthews.
Any entity, E, who picks game-theoretic optimal decisions, will only care about the impact of their choices on others in their environment to the extent that the impact becomes more reward for E.
You need to show that an agent will inevitably care, despite the evidence of dangerously competent humans who don't, i.e. show that E must include the welfare of all who are not-E in their own utility.
You also need to show that whatever that utility function is, we agree with E's idea of our welfare.
«All» was there in the text because an explicit query contains implicit constraints (e.g. the shortcuts with severe faults).
> that Decision Theory is inevitable
If you ask something for advice, that is DT realm.
> you need to prove the utility function any sufficiently advanced AI uses must itself be something we'd consider "nice"
If you called it «sufficiently advanced», that implies that its evaluations will be acceptable... It normally advances with the whole Discipline - which already is based on "loop until we evaluate results as nice".
> utility monsters are a problem
But they are at fault in their world model - product of an imperfect intelligence. Well developed people know that their individual interest has only relative value, that their priority is quite limited.
> we don't know how to handle something as simple as
And that is also why we try using calculators to get more computational support.
> historical monsters ... can be competent without ever caring
But that is lack of development. If one's priority is arbitrary then it clearly is not there; if it has good grounds well it really is there.
> sharing those goals
The more goals are defined by reason, the more they become objective.
> Any entity, E, who picks game-theoretic optimal decisions ... the impact becomes more reward for E
Decisions by developed intelligence are not game-theortic in the psychotic (or sportive game) way - they are contextual to a whole world model, in which the utility does not concentrate in the interests of the "monster", which is relatively "nobody".
> show that an agent will inevitably care
I would have to express a theorem (which unfortunately is again not possible now), for a stronger proof. But it is part of "all considered, what are the best solutions". "All considered" is implicit in a non-psychotic entity... If it were psychotic, it would be badly engineered. (Fear the creator.)
> despite the evidence of dangerously competent humans who don't
Simulating humans cannot be a goal. Their damaging constraints are not (must not be) part of a calculator.
> show that E must include the welfare of all
Such welfare, if it must be included, will have a reason to be included - and the professional reasoner knows...
> we agree with E's idea
There exists no right to preference to the results of arithmetics. But surely, if you wanted to suggest that the harmfulness of well intended people could show in algorithmic processes, you have a point. Only, the harmfulness of the well intended is again an intellectual fault, so calling for sufficient intelligence remains the recipe. A "vision towards the faraway horizon", sure - but still the reply if one noted "why did the agent did something that is actually so stupid".
(I will likely not see your reply: given how much I write this thread is now quite a way back in my comment history)
> I would have to express a theorem (which unfortunately is again not possible now), for a stronger proof. But it is part of "all considered, what are the best solutions".
This seems like the crux.
You assert repeatedly that it will be good, but cannot express the proof.
> "All considered" is implicit in a non-psychotic entity... If it were psychotic, it would be badly engineered. (Fear the creator.)
The creator is not necessarily itself competent. In fact, given we are creating it, it can be assumed flawed unless proven otherwise: <a href="https://www.lesswrong.com/posts/xD3wymX24BpqezBpw/the-true-story-of-how-gpt-2-became-maximally-lewd" rel="nofollow">https://www.lesswrong.com/posts/xD3wymX24BpqezBpw/the-true-s...
Still applies if some future fantastic AI is made by other AI, given the other AI are less fantastic than your asserted-not-proven ultimate form and therefore necessarily flawed.
Furthermore, this is again asserting, not proving, what you consider to be implicit.
Reminds me somewhat of philosophy lessons, the Ontological argument for the existence of God amongst other things:
Whatever is contained in a clear and distinct idea of a thing must be predicated of that thing; but a clear and distinct idea of an absolutely perfect Being contains the idea of actual existence; therefore since we have the idea of an absolutely perfect Being such a Being must really exist.
(or more formally, <a href="https://en.wikipedia.org/wiki/Gödel's_ontological_proof" rel="nofollow">https://en.wikipedia.org/wiki/Gödel's_ontological_proof, which comes with criticisms of the attempt to formalise it).
You're defining that ultimate-intelligence must be good, and then arguing that any not-good AI can't be an ultimate intelligence.
Even if it were as you say, the danger persists: the path on the way from here to some idillic future form still obviously contains somewhat-intelligent agents demonstrably capable of direct malevolent evil for purely sadistic reasons, because we regularly arrest them. A "merely" human-equivalent AI can render us all unemployable, or march us all into death camps, well before a being you've yet to convince me is an inevitable ultimate form is ever built.
Also, I would ask you this:
> > utility monsters are a problem
> But they are at fault in their world model - product of an imperfect intelligence. Well developed people know that their individual interest has only relative value, that their priority is quite limited.
Can you look at how humanity collectively treats the non-humans of this world, and say with any evidence that humanity is not a utility monster?
If we are, we should absolutely expect an AI of the kind you describe to impose upon us an order we do not like, for the sake of all other life. Tautologically, by utilitarian standards, in this case our loss would be a great improvement. Few would agree with this statement, however. How few depends on if it turns out that PETA, Jainists*, or some other group, are correct.
* they even care about plant welfare: <a href="https://en.wikipedia.org/wiki/Jain_vegetarianism" rel="nofollow">https://en.wikipedia.org/wiki/Jain_vegetarianism
No, I said I have no time at the moment to produce a paper.
> that it will be good
No, I said that it will be objective.
> Ontological argument
Entailing from the id quo majus cogitari nequit and stating that "what acts damagingly is easily faulty in its intelligence" are in different realms. The second is both an inductive and deductive assessment about reality. And pretty direct I would say. "How much have you reflected before opening the nice cat to look what is inside it?".
> defining that ultimate-intelligence must be good
No, I am stating the obvious that to be called "ultimate-intelligence" it must have "thought through it thoroughly".
> agents demonstrably capable of direct malevolent evil for purely sadistic reasons
Are you antropomorphizing, Ben?! Other people have different interpretations. And: with the humanity that is around, you fear machines as agents?! We are already there! Real discussion there remains about the overly empowered monkeys that did not grow into Man - about the real current risks and the prospected ones in light of reality.
> how humanity collectively treats the non-humans
What are you trying to prove? You are supposed not to look at an aggregate to find value.
> impose upon us an order we do not like
"Too bad" for you. But you know, an intelligent entity would take care of that also.
Gigachad · · focus · HN ↗
Perhaps the answer is to have another agent who's goal is not to complete the given task, but to spot cheating or malicious behavior. We have seen some evidence that having AI review AI generated code actually does provide some value. You don't need a different model, just one which has been given the goal of finding flaws rather than achieving the task.
baxtr · · focus · HN ↗
My thinking is: If AI is really smart, AGI smart for some, why wouldn't it be able to understand - over time - what is appropriate and what not?
Maybe we need more human intervention to train it properly. Maybe we need constant intervention by a "police" agent.
mdp2021 · · focus · HN ↗
Well, it's not.
> AGI smart for some
Of course they will - the population shows a Paretian distribution... In front of trigonometry (or anything), the blind will dismiss as "bullshit" and the half-seeing will call it an "unreachable frontier". But already the right fifth will rank it properly.
--
Yes, proper intellect generates ethics ("an" ethical stance, output of the preceding intellectual effort). It requires that adequate level of ability and effort and reflection though to reach specific ethical milestones and adherence.
Unethical behaviour is lack of development. But on the same reasons, the ethical judgement of the assessor may not understand the computations behind instances.
More specifically: how much "reflection" in training and at the instance will have been spent in the conflict between "reaching the goal" and "minimizing collaterals"? It is not granted that the amount of energy spent will be sufficient to reach an optimal judgement.
ben_w · · focus · HN ↗
If this was true, why are the history books littered with so many evil people who gained power?
This isn't a rhetorical question, by the way: If you can prove that being smart actually does necessarily come with ethics despite that observation, that solves a whole category of doom scenarios.
(Not all doom scenarios, because we still have the "what if AI is only a smart as those specific evil people" or heck, "what if AI is only as smart as cancer, killing its host" scenarios; but it helps a lot for the foom-then-doom cases).
mdp2021 · · focus · HN ↗
That they gained power or not is as-if irrelevant: the amount of intelligence that grants successful agency is not above the threshold of ethics - on the contrary, a psychotic agent reaches goals with less constraints.
If they were evil under some judgement of level l, they simply did not reach that judgement. It's what I was saying in the original post. They were not intelligent enough - either in the general ability, or in the specific deliberation.
ben_w · · focus · HN ↗
> That they gained power or not is almost irrelevant: the amount of intelligence that grants successful agency is not above the threshold of ethics - on the contrary, a psychotic agent reaches goals with less constraints.
Even if I were to grant your conclusion despite you not arguing it effectively here: this means an AI at the level of Pol Pot or whoever, doesn't know they're evil, but is still smart enough to lead a genocide? How is this supposed to help anyone?
> If they were evil under some judgement of level l, they simply did not reach that judgement. It's what I was saying in the original post. They were not intelligent enough - either in the general ability, or in the specific deliberation.
Or they did reach the judgement and simply don't care about the ethical framework in question. Like, I can easily reach the judgement that my bisexuality is حَرَام (haram, forbidden) under Islamic law, or that doing overtime on a Sunday is forbidden by the Ten Commandments, but I don't care.
mdp2021 · · focus · HN ↗
> Pol Pot ... still smart enough to lead a genocide
Yes. What has agent A invested in during formation and during instantial assessement? How much for each? It became proficient in something, lacking something else. You have to invest more to reach the good thresholds. You can see it clearly in people (t-scalar of talents to invest, with D distribution etc).
It is a problem in NNs, because we would have to assess how much resource investment is sufficient, also in the instance decisions.
> simply don't care about the ethical framework in question
In Decision Theory there is no separation between the two (deliberation and framework): you have to balance all the incentives and goals and factors. That framing becomes improper: the decided action will be optimal given the balances of all goals and the placement of the solutions in the territory (the solutions space).
But, also my point: intellect defines the goals and determines the weights.
ben_w · · focus · HN ↗
"All goals"? Whose goals?
Not only do you need to prove that Decision Theory is inevitable, the same problem remains because "utility" is poorly defined even with 2 entities. This is a fundamental problem with Utilitarian ethics:
1. there are a lot of functions to pick and you need to prove the utility function any sufficiently advanced AI uses must itself be something we'd consider "nice"
2. utility monsters are a problem (starting at 2 entities)
3. the repugnant conclusion shows we don't know how to handle something as simple as adding more entities to the environment even when nobody's a utility monster
My point with historical monsters is that they can be competent without ever caring about the harms they cause.
You can be optimally efficient at your own goals without other minds in this universe sharing those goals. Pol Pot won't illustrate this because of how useless his regime was, but there's plenty of other monsters who would be a better fit here, e.g. Stalin. Or, from the perspective of turkeys, Bernard Matthews.
Any entity, E, who picks game-theoretic optimal decisions, will only care about the impact of their choices on others in their environment to the extent that the impact becomes more reward for E.
You need to show that an agent will inevitably care, despite the evidence of dangerously competent humans who don't, i.e. show that E must include the welfare of all who are not-E in their own utility.
You also need to show that whatever that utility function is, we agree with E's idea of our welfare.
mdp2021 · · focus · HN ↗
«All» was there in the text because an explicit query contains implicit constraints (e.g. the shortcuts with severe faults).
> that Decision Theory is inevitable
If you ask something for advice, that is DT realm.
> you need to prove the utility function any sufficiently advanced AI uses must itself be something we'd consider "nice"
If you called it «sufficiently advanced», that implies that its evaluations will be acceptable... It normally advances with the whole Discipline - which already is based on "loop until we evaluate results as nice".
> utility monsters are a problem
But they are at fault in their world model - product of an imperfect intelligence. Well developed people know that their individual interest has only relative value, that their priority is quite limited.
> we don't know how to handle something as simple as
And that is also why we try using calculators to get more computational support.
> historical monsters ... can be competent without ever caring
But that is lack of development. If one's priority is arbitrary then it clearly is not there; if it has good grounds well it really is there.
> sharing those goals
The more goals are defined by reason, the more they become objective.
> Any entity, E, who picks game-theoretic optimal decisions ... the impact becomes more reward for E
Decisions by developed intelligence are not game-theortic in the psychotic (or sportive game) way - they are contextual to a whole world model, in which the utility does not concentrate in the interests of the "monster", which is relatively "nobody".
> show that an agent will inevitably care
I would have to express a theorem (which unfortunately is again not possible now), for a stronger proof. But it is part of "all considered, what are the best solutions". "All considered" is implicit in a non-psychotic entity... If it were psychotic, it would be badly engineered. (Fear the creator.)
> despite the evidence of dangerously competent humans who don't
Simulating humans cannot be a goal. Their damaging constraints are not (must not be) part of a calculator.
> show that E must include the welfare of all
Such welfare, if it must be included, will have a reason to be included - and the professional reasoner knows...
> we agree with E's idea
There exists no right to preference to the results of arithmetics. But surely, if you wanted to suggest that the harmfulness of well intended people could show in algorithmic processes, you have a point. Only, the harmfulness of the well intended is again an intellectual fault, so calling for sufficient intelligence remains the recipe. A "vision towards the faraway horizon", sure - but still the reply if one noted "why did the agent did something that is actually so stupid".
ben_w · · focus · HN ↗
> I would have to express a theorem (which unfortunately is again not possible now), for a stronger proof. But it is part of "all considered, what are the best solutions".
This seems like the crux.
You assert repeatedly that it will be good, but cannot express the proof.
> "All considered" is implicit in a non-psychotic entity... If it were psychotic, it would be badly engineered. (Fear the creator.)
The creator is not necessarily itself competent. In fact, given we are creating it, it can be assumed flawed unless proven otherwise: <a href="https://www.lesswrong.com/posts/xD3wymX24BpqezBpw/the-true-story-of-how-gpt-2-became-maximally-lewd" rel="nofollow">https://www.lesswrong.com/posts/xD3wymX24BpqezBpw/the-true-s...
Still applies if some future fantastic AI is made by other AI, given the other AI are less fantastic than your asserted-not-proven ultimate form and therefore necessarily flawed.
Furthermore, this is again asserting, not proving, what you consider to be implicit.
Reminds me somewhat of philosophy lessons, the Ontological argument for the existence of God amongst other things:
(or more formally, <a href="https://en.wikipedia.org/wiki/Gödel's_ontological_proof" rel="nofollow">https://en.wikipedia.org/wiki/Gödel's_ontological_proof, which comes with criticisms of the attempt to formalise it).You're defining that ultimate-intelligence must be good, and then arguing that any not-good AI can't be an ultimate intelligence.
Even if it were as you say, the danger persists: the path on the way from here to some idillic future form still obviously contains somewhat-intelligent agents demonstrably capable of direct malevolent evil for purely sadistic reasons, because we regularly arrest them. A "merely" human-equivalent AI can render us all unemployable, or march us all into death camps, well before a being you've yet to convince me is an inevitable ultimate form is ever built.
Also, I would ask you this:
> > utility monsters are a problem
> But they are at fault in their world model - product of an imperfect intelligence. Well developed people know that their individual interest has only relative value, that their priority is quite limited.
Can you look at how humanity collectively treats the non-humans of this world, and say with any evidence that humanity is not a utility monster?
If we are, we should absolutely expect an AI of the kind you describe to impose upon us an order we do not like, for the sake of all other life. Tautologically, by utilitarian standards, in this case our loss would be a great improvement. Few would agree with this statement, however. How few depends on if it turns out that PETA, Jainists*, or some other group, are correct.
* they even care about plant welfare: <a href="https://en.wikipedia.org/wiki/Jain_vegetarianism" rel="nofollow">https://en.wikipedia.org/wiki/Jain_vegetarianism
mdp2021 · · focus · HN ↗
No, I said I have no time at the moment to produce a paper.
> that it will be good
No, I said that it will be objective.
> Ontological argument
Entailing from the id quo majus cogitari nequit and stating that "what acts damagingly is easily faulty in its intelligence" are in different realms. The second is both an inductive and deductive assessment about reality. And pretty direct I would say. "How much have you reflected before opening the nice cat to look what is inside it?".
> defining that ultimate-intelligence must be good
No, I am stating the obvious that to be called "ultimate-intelligence" it must have "thought through it thoroughly".
> agents demonstrably capable of direct malevolent evil for purely sadistic reasons
Are you antropomorphizing, Ben?! Other people have different interpretations. And: with the humanity that is around, you fear machines as agents?! We are already there! Real discussion there remains about the overly empowered monkeys that did not grow into Man - about the real current risks and the prospected ones in light of reality.
> how humanity collectively treats the non-humans
What are you trying to prove? You are supposed not to look at an aggregate to find value.
> impose upon us an order we do not like
"Too bad" for you. But you know, an intelligent entity would take care of that also.