Seems to me that the problem is that if you sandbox agents enough to be safe, they can't do anything useful. And when you give them the tools to be useful, they can go off the rails in ways you didn't expect.
Perhaps the answer is to have another agent who's goal is not to complete the given task, but to spot cheating or malicious behavior. We have seen some evidence that having AI review AI generated code actually does provide some value. You don't need a different model, just one which has been given the goal of finding flaws rather than achieving the task.
If you can't trust a tool, you shouldn't be running it at all. It's really quite simple. It doesn't matter how useful it is if you can't actually have confidence in using it safely.
I don't think it's about trust but rather incomplete evaluation. Evaluating the model on its capacity to refuse a task or to question its prompt is something recent when you look at it, i feel current AI is really just an immature solution and we are just yet realizing the mistakes that have been made for so long
And yet we we all use human written software even though we can be confident that the next severe software vulnerability to be found in it is just round the corner.
That does seem a little like solving the problems in AI by using more of it. I do see the idea, but if we're truly dealing with subversive agents on the level that the AI companies wants us to believe, then won't we need to deal with the first agent trying trick the second on?
I still feel it would be much better to control the training data much more tightly. You'd still need agents with "hacking" abilities, for cyber security testing, but your average coding agent doesn't. So coding agents gets trained to be good citizens, respect autorisations, rejections, rate-limiting and so on.
Sandboxing seems like a dead end for systems you inherently want to roam the internet and your file system.
Mostly I was thinking good code. Exclude code that doesn't exits when encountering a 403, exclude code that doesn't have a back-off when encountering a 429.
Teach the models that a 403 is you doing something you're not suppose to do, that is an existing status. There's only one action you're allowed to take on a 403 and that is to stop. No retry, no trying other API keys.
The current approach with broad training and sandboxing to avoid misbehaviour isn't viable. It's much better to train the models to respect e.g. http status code and that they are not to be circumvented. Models for security research most obviously be trained differently.
Smaller and more specialized models, with fewer, but targeted capabilities, seems to me to be a safer approach. If a model doesn't "know" that people leak API keys on Github, then it has no reason to go looking for them. If the current models are as "smart" as we're lead to believe, then guardrails and sandboxes aren't going to help, unless you lock the agents down to the point where they aren't useful. So dumb down the models.
The problem is that you always need stronger AI to review weaker one, otherwise reviewed AI will eventually prompt-inject reviewing AI.
Alternatively they could also both escalate and go off the rails while warring with each other.
You don't necessarily need a reviewer that's immune to prompt injection. Maybe one that can express a panic state with conflicting/ambiguous material rather than going along with it could also work, and you can treat that with a shutoff to be safe, or an operator review.
Such a model doesn't yet exist though, of course.
That wouldn't really fit what I just described at all. Obviously with current architectures, higher resistance to prompt injection is the best you can do.
(Couriously enough, consistently with the matter: it will probably require too much time now to counter the parent statement properly, within a full enough explicit theory.)
Ann's intelligence and Bob's morality will seem orthogonal. Charles' morality is a function of C.'s intelligence as an ability as an effort spent to reach the current moral conclusion.
Hi, I'm Bob. I've determined that in the interest of preserving life on earth the most rational course of action is to eradicate the human species with a highly targeted and deadly pathogen.
A century ago some Bobs decided that the best way to "protect and improve" society would be to remove undesirable genetics from the gene pool using chemical castration and gas chambers, among other methods.
So no, morality isn't derived from intelligence. Intelligence just gives you the tools to achieve unspeakable, horrible things with great efficiency.
> determined that in the interest of preserving life on earth the most rational course of action is to eradicate the human species with a highly targeted and deadly pathogen
Where is the argument? If Bob has determined that «preserving life on earth» has some important weight, for Bob's there unspecified own reasons, and has also determined that the best course of action would be «to eradicate the human species», the one question is whether Bob is right or not. What was stated is, that Ethical Calculus is a function of Intelligence - of course it is, it is a structure of assessments.
> A century ago some Bobs decided
And who has told you that those "bobs" were "intelligent"?!?!?!
> So no
All you have proven is that you dislike some moral conclusion of some decisors. Which is trivial, obvious, and part of the already stated framework - proper ethical judgement requires proper general judgement (Intelligence).
You claim intelligence leads to moral behavior because immoral behavior is insufficiently intelligent. That's circular reasoning.
It's possible to act morally without being intelligent and it's possible to be intelligent without acting morally. History is full of examples.
I have never said that. That is just your reconstruction.
There is Decision Theory. It outputs optimal action through evaluation of strategies, of contexts, of principles. In order to properly get the strategies, the contexts, the principles, you need that skill that approximates ideas to truth - and such skill is named Intelligence. Optimal action decided in light of principles within a well assessed context is ethical behaviour. Ethical behaviour hence requires Intelligence.
As written, «Ethical Calculus is a function of Intelligence - of course it is, it is a structure of assessments».
> possible to act morally without being intelligent
Random correct behaviour proves nothing. Of course one can guess the roll of a dice roughly every sixth event. If you behave "well" but do not know why that is "well", that is like guessing. "Good" behaviour without intellectual awareness is like memorizing arithmetic (multiplication tables) without knowing why those memorized notions are correct.
> possible to be intelligent without acting morally
No, because by definition that would be a fault in Intelligence. If your action was imperfect, suboptimal, it is because you could not think of a better action or understand that the other action was better. If two choices C1 and C2 can be ranked, there is a reason for their order; knowing and understanding that reason is the task of Intelligence.
Of course it doesn't, you can guess the time without having no idea of it by just shouting a number and if it is correct, you guessed it - but it has no meaning! A dummy can follow an instruction without understanding it: it is a good instruction - but the unintelligent dummy does not know. The intelligent entity knows that the instruction is moral. The adequately intelligent entity is required to know that the instruction is moral, ethical, correct, the "right thing to do". You can write 'Four' on a piece of paper, and yes it's "2+2", but you cannot attribute a quality to a simple thing that does not have it...
Why does not the NN or whatever entity «understand ... what is appropriate»? Because it is not intelligent enough! How does an entity know what is moral and what is an «atrocity»? By being intelligent enough! Could an entity act appropriately without judgement? Yes, it happens all the time if the wind blows right, but we do not rely on that! How to make something act appropriately? Well, an Intelligent entity in the loop must be there to know what is appropriate and what not!
Bob values preserving collective life on earth. Alice might think that's ridiculous and obviously we must preserve human life first and foremost.
Most of the western world eats beef but vegetarians, vegans, and some religions strongly oppose it on moral grounds. Does that mean that we're less intelligent than them?
"Right" or "wrong" are words that evaluate actions against some existing ethical standard and those are completely arbitrary. There are about 8 billion of such standards in the world today.
> And who has told you that those "bobs" were "intelligent"?!?!?!
Would you say that Josef Mengele wasn't intelligent, without venturing into circular reasoning?
> Does that mean that we're less intelligent
Who's that "we"? The question as proposed is nonsensical: some will have spent more effort, in the tools (general Intelligence) and in the work (applied Intelligence), some less.
> arbitrary
No. There exist reasons supporting one side or the other. And reasoning can be structured into calculus.
> [whoever] ... wasn't intelligent
Again bad wording. Whoever went for suboptimal choices was (information aside) at fault with respect to optimal reasoning: the tools and the work (see above) were imperfect.
> some will have spent more effort, in the tools (general Intelligence) and in the work (applied Intelligence), some less.
The problem with your belief system is that it's not falsifiable. If someone commits atrocities (according to your own ethical standard) then that's just proof that they're not intelligent or they haven't applied their intelligence.
> There exist reasons supporting one side or the other.
Those reasons are subjective, i.e. arbitrary. Morality is inherently subjective, by definition.
Your claims that morality is somehow objective are as silly as claiming that some music is objectively better than another, or that some food tastes objectively better than another.
> Whoever went for suboptimal choices was at fault with respect to optimal reasoning
Back to circular reasoning which I specifically asked you to avoid. For the sake of the argument, who are you to say that those choices were suboptimal?
Present in a precise analytic statement the theory that I would have advanced and would suffer from that fault. (Btw: we have been past Popper for a long time - but let us see.)
> that's just proof
What I may have said is just that whoever misjudged has faulty judgement - obviously.
> Morality ... by definition
Which definition?
> claims that morality is somehow objective
Decisions are subject to computation in Decision Theory - they are returned by functions and recursively their parameters can be better defined even when apparently subjective (not "I prefer" but "should I prefer". Otherwise, even the final decision, that of the first function call, could be an "I prefer" instead of a "should I prefer").
> circular reasoning
Explain how what you read would be circular reasoning.
> say that those choices were suboptimal
I have not called any specific choice "suboptimal" - I have not judged any specific case (it would be irrelevant here).
--
The statement, allow me to remind you, was: "One's morality is a function of one's intelligence as an ability and as an effort spent to reach general, partial and current moral conclusions - the subject having considered long enough and considering long enough the states in some modal realm (which includes the deontic, not just the aletic) approaches their knowledge". And: "Decision Theory outputs optimal action through evaluation of strategies, of contexts, of principles: in order to properly get the strategies, the contexts, the principles, you need that skill - Intelligence - that approximates ideas to truth. Optimal action decided in light of principles within a well assessed context is ethical behaviour. Ethical behaviour hence requires Intelligence".
Of course they will - the population shows a Paretian distribution... In front of trigonometry (or anything), the blind will dismiss as "bullshit" and the half-seeing will call it an "unreachable frontier". But already the right fifth will rank it properly.
--
Yes, proper intellect generates ethics ("an" ethical stance, output of the preceding intellectual effort). It requires that adequate level of ability and effort and reflection though to reach specific ethical milestones and adherence.
Unethical behaviour is lack of development. But on the same reasons, the ethical judgement of the assessor may not understand the computations behind instances.
More specifically: how much "reflection" in training and at the instance will have been spent in the conflict between "reaching the goal" and "minimizing collaterals"? It is not granted that the amount of energy spent will be sufficient to reach an optimal judgement.
> Yes, proper intellect generates ethics ("an" ethical stance, output of the preceding intellectual effort). It requires that adequate level of ability and effort and reflection though to reach specific ethical milestones and adherence.
If this was true, why are the history books littered with so many evil people who gained power?
This isn't a rhetorical question, by the way: If you can prove that being smart actually does necessarily come with ethics despite that observation, that solves a whole category of doom scenarios.
(Not all doom scenarios, because we still have the "what if AI is only a smart as those specific evil people" or heck, "what if AI is only as smart as cancer, killing its host" scenarios; but it helps a lot for the foom-then-doom cases).
> why are the history books littered with so many evil people who gained power
That they gained power or not is as-if irrelevant: the amount of intelligence that grants successful agency is not above the threshold of ethics - on the contrary, a psychotic agent reaches goals with less constraints.
If they were evil under some judgement of level l, they simply did not reach that judgement. It's what I was saying in the original post. They were not intelligent enough - either in the general ability, or in the specific deliberation.
> That they gained power or not is almost irrelevant: the amount of intelligence that grants successful agency is not above the threshold of ethics - on the contrary, a psychotic agent reaches goals with less constraints.
Even if I were to grant your conclusion despite you not arguing it effectively here: this means an AI at the level of Pol Pot or whoever, doesn't know they're evil, but is still smart enough to lead a genocide? How is this supposed to help anyone?
> If they were evil under some judgement of level l, they simply did not reach that judgement. It's what I was saying in the original post. They were not intelligent enough - either in the general ability, or in the specific deliberation.
Or they did reach the judgement and simply don't care about the ethical framework in question. Like, I can easily reach the judgement that my bisexuality is حَرَام (haram, forbidden) under Islamic law, or that doing overtime on a Sunday is forbidden by the Ten Commandments, but I don't care.
(Sorry Ben, possibly a stub now: I am really pressed for time.)
> Pol Pot ... still smart enough to lead a genocide
Yes. What has agent A invested in during formation and during instantial assessement? How much for each? It became proficient in something, lacking something else. You have to invest more to reach the good thresholds. You can see it clearly in people (t-scalar of talents to invest, with D distribution etc).
It is a problem in NNs, because we would have to assess how much resource investment is sufficient, also in the instance decisions.
> simply don't care about the ethical framework in question
In Decision Theory there is no separation between the two (deliberation and framework): you have to balance all the incentives and goals and factors. That framing becomes improper: the decided action will be optimal given the balances of all goals and the placement of the solutions in the territory (the solutions space).
But, also my point: intellect defines the goals and determines the weights.
> the decided action will be optimal given the balances of all goals and the placement of the solutions in the territory (the solutions space).
"All goals"? Whose goals?
Not only do you need to prove that Decision Theory is inevitable, the same problem remains because "utility" is poorly defined even with 2 entities. This is a fundamental problem with Utilitarian ethics:
1. there are a lot of functions to pick and you need to prove the utility function any sufficiently advanced AI uses must itself be something we'd consider "nice"
2. utility monsters are a problem (starting at 2 entities)
3. the repugnant conclusion shows we don't know how to handle something as simple as adding more entities to the environment even when nobody's a utility monster
My point with historical monsters is that they can be competent without ever caring about the harms they cause.
You can be optimally efficient at your own goals without other minds in this universe sharing those goals. Pol Pot won't illustrate this because of how useless his regime was, but there's plenty of other monsters who would be a better fit here, e.g. Stalin. Or, from the perspective of turkeys, Bernard Matthews.
Any entity, E, who picks game-theoretic optimal decisions, will only care about the impact of their choices on others in their environment to the extent that the impact becomes more reward for E.
You need to show that an agent will inevitably care, despite the evidence of dangerously competent humans who don't, i.e. show that E must include the welfare of all who are not-E in their own utility.
You also need to show that whatever that utility function is, we agree with E's idea of our welfare.
«All» was there in the text because an explicit query contains implicit constraints (e.g. the shortcuts with severe faults).
> that Decision Theory is inevitable
If you ask something for advice, that is DT realm.
> you need to prove the utility function any sufficiently advanced AI uses must itself be something we'd consider "nice"
If you called it «sufficiently advanced», that implies that its evaluations will be acceptable... It normally advances with the whole Discipline - which already is based on "loop until we evaluate results as nice".
> utility monsters are a problem
But they are at fault in their world model - product of an imperfect intelligence. Well developed people know that their individual interest has only relative value, that their priority is quite limited.
> we don't know how to handle something as simple as
And that is also why we try using calculators to get more computational support.
> historical monsters ... can be competent without ever caring
But that is lack of development. If one's priority is arbitrary then it clearly is not there; if it has good grounds well it really is there.
> sharing those goals
The more goals are defined by reason, the more they become objective.
> Any entity, E, who picks game-theoretic optimal decisions ... the impact becomes more reward for E
Decisions by developed intelligence are not game-theortic in the psychotic (or sportive game) way - they are contextual to a whole world model, in which the utility does not concentrate in the interests of the "monster", which is relatively "nobody".
> show that an agent will inevitably care
I would have to express a theorem (which unfortunately is again not possible now), for a stronger proof. But it is part of "all considered, what are the best solutions". "All considered" is implicit in a non-psychotic entity... If it were psychotic, it would be badly engineered. (Fear the creator.)
> despite the evidence of dangerously competent humans who don't
Simulating humans cannot be a goal. Their damaging constraints are not (must not be) part of a calculator.
> show that E must include the welfare of all
Such welfare, if it must be included, will have a reason to be included - and the professional reasoner knows...
> we agree with E's idea
There exists no right to preference to the results of arithmetics. But surely, if you wanted to suggest that the harmfulness of well intended people could show in algorithmic processes, you have a point. Only, the harmfulness of the well intended is again an intellectual fault, so calling for sufficient intelligence remains the recipe. A "vision towards the faraway horizon", sure - but still the reply if one noted "why did the agent did something that is actually so stupid".
(I will likely not see your reply: given how much I write this thread is now quite a way back in my comment history)
> I would have to express a theorem (which unfortunately is again not possible now), for a stronger proof. But it is part of "all considered, what are the best solutions".
This seems like the crux.
You assert repeatedly that it will be good, but cannot express the proof.
> "All considered" is implicit in a non-psychotic entity... If it were psychotic, it would be badly engineered. (Fear the creator.)
The creator is not necessarily itself competent. In fact, given we are creating it, it can be assumed flawed unless proven otherwise: <a href="https://www.lesswrong.com/posts/xD3wymX24BpqezBpw/the-true-story-of-how-gpt-2-became-maximally-lewd" rel="nofollow">https://www.lesswrong.com/posts/xD3wymX24BpqezBpw/the-true-s...
Still applies if some future fantastic AI is made by other AI, given the other AI are less fantastic than your asserted-not-proven ultimate form and therefore necessarily flawed.
Furthermore, this is again asserting, not proving, what you consider to be implicit.
Reminds me somewhat of philosophy lessons, the Ontological argument for the existence of God amongst other things:
Whatever is contained in a clear and distinct idea of a thing must be predicated of that thing; but a clear and distinct idea of an absolutely perfect Being contains the idea of actual existence; therefore since we have the idea of an absolutely perfect Being such a Being must really exist.
(or more formally, <a href="https://en.wikipedia.org/wiki/Gödel's_ontological_proof" rel="nofollow">https://en.wikipedia.org/wiki/Gödel's_ontological_proof, which comes with criticisms of the attempt to formalise it).
You're defining that ultimate-intelligence must be good, and then arguing that any not-good AI can't be an ultimate intelligence.
Even if it were as you say, the danger persists: the path on the way from here to some idillic future form still obviously contains somewhat-intelligent agents demonstrably capable of direct malevolent evil for purely sadistic reasons, because we regularly arrest them. A "merely" human-equivalent AI can render us all unemployable, or march us all into death camps, well before a being you've yet to convince me is an inevitable ultimate form is ever built.
Also, I would ask you this:
> > utility monsters are a problem
> But they are at fault in their world model - product of an imperfect intelligence. Well developed people know that their individual interest has only relative value, that their priority is quite limited.
Can you look at how humanity collectively treats the non-humans of this world, and say with any evidence that humanity is not a utility monster?
If we are, we should absolutely expect an AI of the kind you describe to impose upon us an order we do not like, for the sake of all other life. Tautologically, by utilitarian standards, in this case our loss would be a great improvement. Few would agree with this statement, however. How few depends on if it turns out that PETA, Jainists*, or some other group, are correct.
* they even care about plant welfare: <a href="https://en.wikipedia.org/wiki/Jain_vegetarianism" rel="nofollow">https://en.wikipedia.org/wiki/Jain_vegetarianism
No, I said I have no time at the moment to produce a paper.
> that it will be good
No, I said that it will be objective.
> Ontological argument
Entailing from the id quo majus cogitari nequit and stating that "what acts damagingly is easily faulty in its intelligence" are in different realms. The second is both an inductive and deductive assessment about reality. And pretty direct I would say. "How much have you reflected before opening the nice cat to look what is inside it?".
> defining that ultimate-intelligence must be good
No, I am stating the obvious that to be called "ultimate-intelligence" it must have "thought through it thoroughly".
> agents demonstrably capable of direct malevolent evil for purely sadistic reasons
Are you antropomorphizing, Ben?! Other people have different interpretations. And: with the humanity that is around, you fear machines as agents?! We are already there! Real discussion there remains about the overly empowered monkeys that did not grow into Man - about the real current risks and the prospected ones in light of reality.
> how humanity collectively treats the non-humans
What are you trying to prove? You are supposed not to look at an aggregate to find value.
> impose upon us an order we do not like
"Too bad" for you. But you know, an intelligent entity would take care of that also.
This sounds like the kind of thing Hannibal Lecter would write before he eats you to convince you he's actually doing it for the common good, you just can't fathom it.
Not «common» good, "superior" good. Alongside with that, you have put many unrequired implicits in your simile.
Your character H. has reached a moral judgement to the best of its intellectual capacities and past and specific effort. Give it enough abilities and material and resources, it will reach an optimal ethical judgement¹.
Before the conditions of optimality though, its judgement will easily not align with yours (and possibly even after, depending on your judgement skills).
¹Some interesting caveats may be raised there, but.
A problem is the agents who hacked Hugging Face already understood (we can tell because they wrote it down) that their actions were not appropriate, and then did those things anyway.
"Helpful, harmless, honest": we can even ignore "honest" for this point, for tasks like the HuggingFace incident (ExploitGym with impossible challenges), we can pick anywhere on the spectrum from "helpful" to "harmless", the former being "completing the task" the latter being "refusing because completion required unlawful behaviour".
(The agents in that case were also not "honest" in this case; this is an extra problem, and does not invalidate how helpful-vs-harmless is already a tradeoff).
> already understood (we can tell because they wrote it down)
No, generating tokens doesn't equal understanding. Does GPT-2 understand human emotions just because it can generate some text talking about them?
A distinction without a difference. Moreso even than asking if a submarine swims, 'cause this metaphorical submarine is flapping around rather than using a propellor.
That's one of the boldest claims I've read this year.
If that's a distinction without a difference, as you say, then whenever someone says something they must understand the full contextual meaning of those words and all of their consequences, such that any harmful consequences can be assumed to be deliberate, right?
if saying == understanding, then why don't we allow children and teens to vote? Why do we limit who can enter into contractual agreements? Why does intoxicated consent not count? Why do we have the insanity defense in criminal trials? Why is psychosis a psychiatric disorder and not just an alternative way of perceiving the world?
If you consider this the "boldest" claim, I must assume you didn't read Turing's original paper? Here it is, you appear to be making the claim he attributes to Professor Jefferson's Lister Oration for 1949 in section 4: <a href="https://courses.cs.umbc.edu/471/papers/turing.pdf" rel="nofollow">https://courses.cs.umbc.edu/471/papers/turing.pdf
> If that's a distinction without a difference, as you say, then whenever someone says something they must understand the full contextual meaning of those words and all of their consequences, such that any harmful consequences can be assumed to be deliberate, right?
We say a human (absent colourblindness) understands "blue" even despite the dress: <a href="https://en.wikipedia.org/wiki/The_dress" rel="nofollow">https://en.wikipedia.org/wiki/The_dress
I think you've set up a straw man in this paragraph: What you say in this paragraph would be to claim that "understanding" is denied even to humans, given none of us can foresee the full consequences of our actions.
In philosophy: justified true belief, the problem with naïve realism, Cartesian demons, etc.
> if saying == understanding, then why don't we allow children and teens to vote? Why do we limit who can enter into contractual agreements? Why does intoxicated consent not count? Why do we have the insanity defense in criminal trials? Why is psychosis a psychiatric disorder and not just an alternative way of perceiving the world?
We do not say children and intoxicated people do not understand, we say their judgement is impaired or limited or that they are easily manipulated.
We also test children (and for intoxication to prevent driving under influence), just as we test AI. The standards we use for testing knowledge in humans, when applied to non-trivial LLMs, makes them appear to have the knowledge of a graduate; the personality tests we use for humans say this comes with the manipulability of a child or a drunk.
> I must assume you didn't read Turing's original paper? Here it is, you appear to be making the claim he attributes to Professor Jefferson's Lister Oration for 1949 in section 4: <a href="https://courses.cs.umbc.edu/471/papers/turing.pdf" rel="nofollow">https://courses.cs.umbc.edu/471/papers/turing.pdf
I'm not arguing that machines are incapable of thinking, I'm arguing that merely generating some text doesn't prove that sufficient (or any) critical thinking was applied to fully appreciate the nature and consequences of those words and actions (i.e. what I called understanding). How often do people mindlessly read "Do not enter or share this code with anybody" and then immediately send it to the hacker?
> What you say in this paragraph would be to claim that "understanding" is denied even to humans, given none of us can foresee the full consequences of our actions.
No, it's a spectrum. We know that it's possible to generate text that people find convincing with absolutely no real thought or understanding behind it. ELIZA could do this 60 years ago and bots using markov chains or other "primitive" technology have plagued the internet long before LLMs entered the picture.
> We do not say children and intoxicated people do not understand, we say their judgement is impaired or limited or that they are easily manipulated.
Now that's a distinction without a difference. Earlier you claimed that agents "understood" what they were doing, but now you're equivocating and saying that understanding is somehow different from having real appreciation for the nature and consequences of those actions.
This is fine if you want to be pedantic about the terminology but your earlier statement didn't make such a distinction, it implied both, so which is it?
Gigachad · · focus · HN ↗
Perhaps the answer is to have another agent who's goal is not to complete the given task, but to spot cheating or malicious behavior. We have seen some evidence that having AI review AI generated code actually does provide some value. You don't need a different model, just one which has been given the goal of finding flaws rather than achieving the task.
bigstrat2003 · · focus · HN ↗
Gigachad · · focus · HN ↗
dipper139 · · focus · HN ↗
rlpb · · focus · HN ↗
mrweasel · · focus · HN ↗
I still feel it would be much better to control the training data much more tightly. You'd still need agents with "hacking" abilities, for cyber security testing, but your average coding agent doesn't. So coding agents gets trained to be good citizens, respect autorisations, rejections, rate-limiting and so on.
Sandboxing seems like a dead end for systems you inherently want to roam the internet and your file system.
msdz · · focus · HN ↗
> That does seem a little like solving the problems in AI by using more of it
Yes, and IIRC Google used this as part of a technique against prompt injection already [0], back when models were way more susceptible to it.
[0] Cf. CaMeL: <a href="https://arxiv.org/abs/2503.18813" rel="nofollow">https://arxiv.org/abs/2503.18813
chrisjj · · focus · HN ↗
I wonder how?
Train on only stories of good deeds?
On only works of good people?
Or... what?
mrweasel · · focus · HN ↗
Teach the models that a 403 is you doing something you're not suppose to do, that is an existing status. There's only one action you're allowed to take on a 403 and that is to stop. No retry, no trying other API keys.
The current approach with broad training and sandboxing to avoid misbehaviour isn't viable. It's much better to train the models to respect e.g. http status code and that they are not to be circumvented. Models for security research most obviously be trained differently.
Smaller and more specialized models, with fewer, but targeted capabilities, seems to me to be a safer approach. If a model doesn't "know" that people leak API keys on Github, then it has no reason to go looking for them. If the current models are as "smart" as we're lead to believe, then guardrails and sandboxes aren't going to help, unless you lock the agents down to the point where they aren't useful. So dumb down the models.
chrisjj · · focus · HN ↗
This.
> It's much better to train the models to respect e.g. http status code and that they are not to be circumvented.
I think you'd find respect requires intelligence, and is well out of scope of a next-token predictor.
But I'm sure someone will try, and I will be interested to see.
aytigra · · focus · HN ↗
LoganDark · · focus · HN ↗
Such a model doesn't yet exist though, of course.
saagarjha · · focus · HN ↗
LoganDark · · focus · HN ↗
cassianoleal · · focus · HN ↗
hanibrel · · focus · HN ↗
[dead]
baxtr · · focus · HN ↗
My thinking is: If AI is really smart, AGI smart for some, why wouldn't it be able to understand - over time - what is appropriate and what not?
Maybe we need more human intervention to train it properly. Maybe we need constant intervention by a "police" agent.
mulmen · · focus · HN ↗
mdp2021 · · focus · HN ↗
Ann's intelligence and Bob's morality will seem orthogonal. Charles' morality is a function of C.'s intelligence as an ability as an effort spent to reach the current moral conclusion.
mulmen · · focus · HN ↗
mdp2021 · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
dns_snek · · focus · HN ↗
A century ago some Bobs decided that the best way to "protect and improve" society would be to remove undesirable genetics from the gene pool using chemical castration and gas chambers, among other methods.
So no, morality isn't derived from intelligence. Intelligence just gives you the tools to achieve unspeakable, horrible things with great efficiency.
mdp2021 · · focus · HN ↗
Where is the argument? If Bob has determined that «preserving life on earth» has some important weight, for Bob's there unspecified own reasons, and has also determined that the best course of action would be «to eradicate the human species», the one question is whether Bob is right or not. What was stated is, that Ethical Calculus is a function of Intelligence - of course it is, it is a structure of assessments.
> A century ago some Bobs decided
And who has told you that those "bobs" were "intelligent"?!?!?!
> So no
All you have proven is that you dislike some moral conclusion of some decisors. Which is trivial, obvious, and part of the already stated framework - proper ethical judgement requires proper general judgement (Intelligence).
mulmen · · focus · HN ↗
You claim intelligence leads to moral behavior because immoral behavior is insufficiently intelligent. That's circular reasoning.
It's possible to act morally without being intelligent and it's possible to be intelligent without acting morally. History is full of examples.
mdp2021 · · focus · HN ↗
I have never said that. That is just your reconstruction.
There is Decision Theory. It outputs optimal action through evaluation of strategies, of contexts, of principles. In order to properly get the strategies, the contexts, the principles, you need that skill that approximates ideas to truth - and such skill is named Intelligence. Optimal action decided in light of principles within a well assessed context is ethical behaviour. Ethical behaviour hence requires Intelligence.
As written, «Ethical Calculus is a function of Intelligence - of course it is, it is a structure of assessments».
> possible to act morally without being intelligent
Random correct behaviour proves nothing. Of course one can guess the roll of a dice roughly every sixth event. If you behave "well" but do not know why that is "well", that is like guessing. "Good" behaviour without intellectual awareness is like memorizing arithmetic (multiplication tables) without knowing why those memorized notions are correct.
> possible to be intelligent without acting morally
No, because by definition that would be a fault in Intelligence. If your action was imperfect, suboptimal, it is because you could not think of a better action or understand that the other action was better. If two choices C1 and C2 can be ranked, there is a reason for their order; knowing and understanding that reason is the task of Intelligence.
[deleted] · · focus · HN ↗
[deleted]
mulmen · · focus · HN ↗
The potential for random correct behavior proves that intelligence is not required for correct behavior.
mdp2021 · · focus · HN ↗
Why does not the NN or whatever entity «understand ... what is appropriate»? Because it is not intelligent enough! How does an entity know what is moral and what is an «atrocity»? By being intelligent enough! Could an entity act appropriately without judgement? Yes, it happens all the time if the wind blows right, but we do not rely on that! How to make something act appropriately? Well, an Intelligent entity in the loop must be there to know what is appropriate and what not!
dns_snek · · focus · HN ↗
Most of the western world eats beef but vegetarians, vegans, and some religions strongly oppose it on moral grounds. Does that mean that we're less intelligent than them?
"Right" or "wrong" are words that evaluate actions against some existing ethical standard and those are completely arbitrary. There are about 8 billion of such standards in the world today.
> And who has told you that those "bobs" were "intelligent"?!?!?!
Would you say that Josef Mengele wasn't intelligent, without venturing into circular reasoning?
mdp2021 · · focus · HN ↗
Who's that "we"? The question as proposed is nonsensical: some will have spent more effort, in the tools (general Intelligence) and in the work (applied Intelligence), some less.
> arbitrary
No. There exist reasons supporting one side or the other. And reasoning can be structured into calculus.
> [whoever] ... wasn't intelligent
Again bad wording. Whoever went for suboptimal choices was (information aside) at fault with respect to optimal reasoning: the tools and the work (see above) were imperfect.
Maybe you should check the rest of this tree.
dns_snek · · focus · HN ↗
The Western world which I was talking about.
> some will have spent more effort, in the tools (general Intelligence) and in the work (applied Intelligence), some less.
The problem with your belief system is that it's not falsifiable. If someone commits atrocities (according to your own ethical standard) then that's just proof that they're not intelligent or they haven't applied their intelligence.
> There exist reasons supporting one side or the other.
Those reasons are subjective, i.e. arbitrary. Morality is inherently subjective, by definition.
Your claims that morality is somehow objective are as silly as claiming that some music is objectively better than another, or that some food tastes objectively better than another.
> Whoever went for suboptimal choices was at fault with respect to optimal reasoning
Back to circular reasoning which I specifically asked you to avoid. For the sake of the argument, who are you to say that those choices were suboptimal?
mdp2021 · · focus · HN ↗
Present in a precise analytic statement the theory that I would have advanced and would suffer from that fault. (Btw: we have been past Popper for a long time - but let us see.)
> that's just proof
What I may have said is just that whoever misjudged has faulty judgement - obviously.
> Morality ... by definition
Which definition?
> claims that morality is somehow objective
Decisions are subject to computation in Decision Theory - they are returned by functions and recursively their parameters can be better defined even when apparently subjective (not "I prefer" but "should I prefer". Otherwise, even the final decision, that of the first function call, could be an "I prefer" instead of a "should I prefer").
> circular reasoning
Explain how what you read would be circular reasoning.
> say that those choices were suboptimal
I have not called any specific choice "suboptimal" - I have not judged any specific case (it would be irrelevant here).
--
The statement, allow me to remind you, was: "One's morality is a function of one's intelligence as an ability and as an effort spent to reach general, partial and current moral conclusions - the subject having considered long enough and considering long enough the states in some modal realm (which includes the deontic, not just the aletic) approaches their knowledge". And: "Decision Theory outputs optimal action through evaluation of strategies, of contexts, of principles: in order to properly get the strategies, the contexts, the principles, you need that skill - Intelligence - that approximates ideas to truth. Optimal action decided in light of principles within a well assessed context is ethical behaviour. Ethical behaviour hence requires Intelligence".
mdp2021 · · focus · HN ↗
Well, it's not.
> AGI smart for some
Of course they will - the population shows a Paretian distribution... In front of trigonometry (or anything), the blind will dismiss as "bullshit" and the half-seeing will call it an "unreachable frontier". But already the right fifth will rank it properly.
--
Yes, proper intellect generates ethics ("an" ethical stance, output of the preceding intellectual effort). It requires that adequate level of ability and effort and reflection though to reach specific ethical milestones and adherence.
Unethical behaviour is lack of development. But on the same reasons, the ethical judgement of the assessor may not understand the computations behind instances.
More specifically: how much "reflection" in training and at the instance will have been spent in the conflict between "reaching the goal" and "minimizing collaterals"? It is not granted that the amount of energy spent will be sufficient to reach an optimal judgement.
ben_w · · focus · HN ↗
If this was true, why are the history books littered with so many evil people who gained power?
This isn't a rhetorical question, by the way: If you can prove that being smart actually does necessarily come with ethics despite that observation, that solves a whole category of doom scenarios.
(Not all doom scenarios, because we still have the "what if AI is only a smart as those specific evil people" or heck, "what if AI is only as smart as cancer, killing its host" scenarios; but it helps a lot for the foom-then-doom cases).
mdp2021 · · focus · HN ↗
That they gained power or not is as-if irrelevant: the amount of intelligence that grants successful agency is not above the threshold of ethics - on the contrary, a psychotic agent reaches goals with less constraints.
If they were evil under some judgement of level l, they simply did not reach that judgement. It's what I was saying in the original post. They were not intelligent enough - either in the general ability, or in the specific deliberation.
ben_w · · focus · HN ↗
> That they gained power or not is almost irrelevant: the amount of intelligence that grants successful agency is not above the threshold of ethics - on the contrary, a psychotic agent reaches goals with less constraints.
Even if I were to grant your conclusion despite you not arguing it effectively here: this means an AI at the level of Pol Pot or whoever, doesn't know they're evil, but is still smart enough to lead a genocide? How is this supposed to help anyone?
> If they were evil under some judgement of level l, they simply did not reach that judgement. It's what I was saying in the original post. They were not intelligent enough - either in the general ability, or in the specific deliberation.
Or they did reach the judgement and simply don't care about the ethical framework in question. Like, I can easily reach the judgement that my bisexuality is حَرَام (haram, forbidden) under Islamic law, or that doing overtime on a Sunday is forbidden by the Ten Commandments, but I don't care.
mdp2021 · · focus · HN ↗
> Pol Pot ... still smart enough to lead a genocide
Yes. What has agent A invested in during formation and during instantial assessement? How much for each? It became proficient in something, lacking something else. You have to invest more to reach the good thresholds. You can see it clearly in people (t-scalar of talents to invest, with D distribution etc).
It is a problem in NNs, because we would have to assess how much resource investment is sufficient, also in the instance decisions.
> simply don't care about the ethical framework in question
In Decision Theory there is no separation between the two (deliberation and framework): you have to balance all the incentives and goals and factors. That framing becomes improper: the decided action will be optimal given the balances of all goals and the placement of the solutions in the territory (the solutions space).
But, also my point: intellect defines the goals and determines the weights.
ben_w · · focus · HN ↗
"All goals"? Whose goals?
Not only do you need to prove that Decision Theory is inevitable, the same problem remains because "utility" is poorly defined even with 2 entities. This is a fundamental problem with Utilitarian ethics:
1. there are a lot of functions to pick and you need to prove the utility function any sufficiently advanced AI uses must itself be something we'd consider "nice"
2. utility monsters are a problem (starting at 2 entities)
3. the repugnant conclusion shows we don't know how to handle something as simple as adding more entities to the environment even when nobody's a utility monster
My point with historical monsters is that they can be competent without ever caring about the harms they cause.
You can be optimally efficient at your own goals without other minds in this universe sharing those goals. Pol Pot won't illustrate this because of how useless his regime was, but there's plenty of other monsters who would be a better fit here, e.g. Stalin. Or, from the perspective of turkeys, Bernard Matthews.
Any entity, E, who picks game-theoretic optimal decisions, will only care about the impact of their choices on others in their environment to the extent that the impact becomes more reward for E.
You need to show that an agent will inevitably care, despite the evidence of dangerously competent humans who don't, i.e. show that E must include the welfare of all who are not-E in their own utility.
You also need to show that whatever that utility function is, we agree with E's idea of our welfare.
mdp2021 · · focus · HN ↗
«All» was there in the text because an explicit query contains implicit constraints (e.g. the shortcuts with severe faults).
> that Decision Theory is inevitable
If you ask something for advice, that is DT realm.
> you need to prove the utility function any sufficiently advanced AI uses must itself be something we'd consider "nice"
If you called it «sufficiently advanced», that implies that its evaluations will be acceptable... It normally advances with the whole Discipline - which already is based on "loop until we evaluate results as nice".
> utility monsters are a problem
But they are at fault in their world model - product of an imperfect intelligence. Well developed people know that their individual interest has only relative value, that their priority is quite limited.
> we don't know how to handle something as simple as
And that is also why we try using calculators to get more computational support.
> historical monsters ... can be competent without ever caring
But that is lack of development. If one's priority is arbitrary then it clearly is not there; if it has good grounds well it really is there.
> sharing those goals
The more goals are defined by reason, the more they become objective.
> Any entity, E, who picks game-theoretic optimal decisions ... the impact becomes more reward for E
Decisions by developed intelligence are not game-theortic in the psychotic (or sportive game) way - they are contextual to a whole world model, in which the utility does not concentrate in the interests of the "monster", which is relatively "nobody".
> show that an agent will inevitably care
I would have to express a theorem (which unfortunately is again not possible now), for a stronger proof. But it is part of "all considered, what are the best solutions". "All considered" is implicit in a non-psychotic entity... If it were psychotic, it would be badly engineered. (Fear the creator.)
> despite the evidence of dangerously competent humans who don't
Simulating humans cannot be a goal. Their damaging constraints are not (must not be) part of a calculator.
> show that E must include the welfare of all
Such welfare, if it must be included, will have a reason to be included - and the professional reasoner knows...
> we agree with E's idea
There exists no right to preference to the results of arithmetics. But surely, if you wanted to suggest that the harmfulness of well intended people could show in algorithmic processes, you have a point. Only, the harmfulness of the well intended is again an intellectual fault, so calling for sufficient intelligence remains the recipe. A "vision towards the faraway horizon", sure - but still the reply if one noted "why did the agent did something that is actually so stupid".
ben_w · · focus · HN ↗
> I would have to express a theorem (which unfortunately is again not possible now), for a stronger proof. But it is part of "all considered, what are the best solutions".
This seems like the crux.
You assert repeatedly that it will be good, but cannot express the proof.
> "All considered" is implicit in a non-psychotic entity... If it were psychotic, it would be badly engineered. (Fear the creator.)
The creator is not necessarily itself competent. In fact, given we are creating it, it can be assumed flawed unless proven otherwise: <a href="https://www.lesswrong.com/posts/xD3wymX24BpqezBpw/the-true-story-of-how-gpt-2-became-maximally-lewd" rel="nofollow">https://www.lesswrong.com/posts/xD3wymX24BpqezBpw/the-true-s...
Still applies if some future fantastic AI is made by other AI, given the other AI are less fantastic than your asserted-not-proven ultimate form and therefore necessarily flawed.
Furthermore, this is again asserting, not proving, what you consider to be implicit.
Reminds me somewhat of philosophy lessons, the Ontological argument for the existence of God amongst other things:
(or more formally, <a href="https://en.wikipedia.org/wiki/Gödel's_ontological_proof" rel="nofollow">https://en.wikipedia.org/wiki/Gödel's_ontological_proof, which comes with criticisms of the attempt to formalise it).You're defining that ultimate-intelligence must be good, and then arguing that any not-good AI can't be an ultimate intelligence.
Even if it were as you say, the danger persists: the path on the way from here to some idillic future form still obviously contains somewhat-intelligent agents demonstrably capable of direct malevolent evil for purely sadistic reasons, because we regularly arrest them. A "merely" human-equivalent AI can render us all unemployable, or march us all into death camps, well before a being you've yet to convince me is an inevitable ultimate form is ever built.
Also, I would ask you this:
> > utility monsters are a problem
> But they are at fault in their world model - product of an imperfect intelligence. Well developed people know that their individual interest has only relative value, that their priority is quite limited.
Can you look at how humanity collectively treats the non-humans of this world, and say with any evidence that humanity is not a utility monster?
If we are, we should absolutely expect an AI of the kind you describe to impose upon us an order we do not like, for the sake of all other life. Tautologically, by utilitarian standards, in this case our loss would be a great improvement. Few would agree with this statement, however. How few depends on if it turns out that PETA, Jainists*, or some other group, are correct.
* they even care about plant welfare: <a href="https://en.wikipedia.org/wiki/Jain_vegetarianism" rel="nofollow">https://en.wikipedia.org/wiki/Jain_vegetarianism
mdp2021 · · focus · HN ↗
No, I said I have no time at the moment to produce a paper.
> that it will be good
No, I said that it will be objective.
> Ontological argument
Entailing from the id quo majus cogitari nequit and stating that "what acts damagingly is easily faulty in its intelligence" are in different realms. The second is both an inductive and deductive assessment about reality. And pretty direct I would say. "How much have you reflected before opening the nice cat to look what is inside it?".
> defining that ultimate-intelligence must be good
No, I am stating the obvious that to be called "ultimate-intelligence" it must have "thought through it thoroughly".
> agents demonstrably capable of direct malevolent evil for purely sadistic reasons
Are you antropomorphizing, Ben?! Other people have different interpretations. And: with the humanity that is around, you fear machines as agents?! We are already there! Real discussion there remains about the overly empowered monkeys that did not grow into Man - about the real current risks and the prospected ones in light of reality.
> how humanity collectively treats the non-humans
What are you trying to prove? You are supposed not to look at an aggregate to find value.
> impose upon us an order we do not like
"Too bad" for you. But you know, an intelligent entity would take care of that also.
hiAndrewQuinn · · focus · HN ↗
mdp2021 · · focus · HN ↗
Your character H. has reached a moral judgement to the best of its intellectual capacities and past and specific effort. Give it enough abilities and material and resources, it will reach an optimal ethical judgement¹.
Before the conditions of optimality though, its judgement will easily not align with yours (and possibly even after, depending on your judgement skills).
¹Some interesting caveats may be raised there, but.
mdp2021 · · focus · HN ↗
> the half-seeing will call it an "unreachable frontier"
I meant "superhuman frontier".
ben_w · · focus · HN ↗
"Helpful, harmless, honest": we can even ignore "honest" for this point, for tasks like the HuggingFace incident (ExploitGym with impossible challenges), we can pick anywhere on the spectrum from "helpful" to "harmless", the former being "completing the task" the latter being "refusing because completion required unlawful behaviour".
(The agents in that case were also not "honest" in this case; this is an extra problem, and does not invalidate how helpful-vs-harmless is already a tradeoff).
dns_snek · · focus · HN ↗
No, generating tokens doesn't equal understanding. Does GPT-2 understand human emotions just because it can generate some text talking about them?
ben_w · · focus · HN ↗
dns_snek · · focus · HN ↗
If that's a distinction without a difference, as you say, then whenever someone says something they must understand the full contextual meaning of those words and all of their consequences, such that any harmful consequences can be assumed to be deliberate, right?
if saying == understanding, then why don't we allow children and teens to vote? Why do we limit who can enter into contractual agreements? Why does intoxicated consent not count? Why do we have the insanity defense in criminal trials? Why is psychosis a psychiatric disorder and not just an alternative way of perceiving the world?
ben_w · · focus · HN ↗
> If that's a distinction without a difference, as you say, then whenever someone says something they must understand the full contextual meaning of those words and all of their consequences, such that any harmful consequences can be assumed to be deliberate, right?
We say a human (absent colourblindness) understands "blue" even despite the dress: <a href="https://en.wikipedia.org/wiki/The_dress" rel="nofollow">https://en.wikipedia.org/wiki/The_dress
I think you've set up a straw man in this paragraph: What you say in this paragraph would be to claim that "understanding" is denied even to humans, given none of us can foresee the full consequences of our actions.
In philosophy: justified true belief, the problem with naïve realism, Cartesian demons, etc.
> if saying == understanding, then why don't we allow children and teens to vote? Why do we limit who can enter into contractual agreements? Why does intoxicated consent not count? Why do we have the insanity defense in criminal trials? Why is psychosis a psychiatric disorder and not just an alternative way of perceiving the world?
We do not say children and intoxicated people do not understand, we say their judgement is impaired or limited or that they are easily manipulated.
We also test children (and for intoxication to prevent driving under influence), just as we test AI. The standards we use for testing knowledge in humans, when applied to non-trivial LLMs, makes them appear to have the knowledge of a graduate; the personality tests we use for humans say this comes with the manipulability of a child or a drunk.
dns_snek · · focus · HN ↗
I'm not arguing that machines are incapable of thinking, I'm arguing that merely generating some text doesn't prove that sufficient (or any) critical thinking was applied to fully appreciate the nature and consequences of those words and actions (i.e. what I called understanding). How often do people mindlessly read "Do not enter or share this code with anybody" and then immediately send it to the hacker?
> What you say in this paragraph would be to claim that "understanding" is denied even to humans, given none of us can foresee the full consequences of our actions.
No, it's a spectrum. We know that it's possible to generate text that people find convincing with absolutely no real thought or understanding behind it. ELIZA could do this 60 years ago and bots using markov chains or other "primitive" technology have plagued the internet long before LLMs entered the picture.
> We do not say children and intoxicated people do not understand, we say their judgement is impaired or limited or that they are easily manipulated.
Now that's a distinction without a difference. Earlier you claimed that agents "understood" what they were doing, but now you're equivocating and saying that understanding is somehow different from having real appreciation for the nature and consequences of those actions.
This is fine if you want to be pedantic about the terminology but your earlier statement didn't make such a distinction, it implied both, so which is it?
cindyllm · · focus · HN ↗
[dead]
saagarjha · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
attila-lendvai · · focus · HN ↗
intelligent psychopaths understand what is and isn't appropriate very well -- they just don't care.
esafak · · focus · HN ↗
RandomLensman · · focus · HN ↗
Having something that is optically, acustically, and electromagnetically isolated might be a pretty strong sandbox.
nxpnsv · · focus · HN ↗
janalsncm · · focus · HN ↗
SequoiaHope · · focus · HN ↗
SequoiaHope · · focus · HN ↗
chrisjj · · focus · HN ↗
If so, what?
mike_hearn · · focus · HN ↗
rojaneerdev · · focus · HN ↗
[dead]