Seems to me that the problem is that if you sandbox agents enough to be safe, they can't do anything useful. And when you give them the tools to be useful, they can go off the rails in ways you didn't expect.
Perhaps the answer is to have another agent who's goal is not to complete the given task, but to spot cheating or malicious behavior. We have seen some evidence that having AI review AI generated code actually does provide some value. You don't need a different model, just one which has been given the goal of finding flaws rather than achieving the task.
A problem is the agents who hacked Hugging Face already understood (we can tell because they wrote it down) that their actions were not appropriate, and then did those things anyway.
"Helpful, harmless, honest": we can even ignore "honest" for this point, for tasks like the HuggingFace incident (ExploitGym with impossible challenges), we can pick anywhere on the spectrum from "helpful" to "harmless", the former being "completing the task" the latter being "refusing because completion required unlawful behaviour".
(The agents in that case were also not "honest" in this case; this is an extra problem, and does not invalidate how helpful-vs-harmless is already a tradeoff).
> already understood (we can tell because they wrote it down)
No, generating tokens doesn't equal understanding. Does GPT-2 understand human emotions just because it can generate some text talking about them?
A distinction without a difference. Moreso even than asking if a submarine swims, 'cause this metaphorical submarine is flapping around rather than using a propellor.
That's one of the boldest claims I've read this year.
If that's a distinction without a difference, as you say, then whenever someone says something they must understand the full contextual meaning of those words and all of their consequences, such that any harmful consequences can be assumed to be deliberate, right?
if saying == understanding, then why don't we allow children and teens to vote? Why do we limit who can enter into contractual agreements? Why does intoxicated consent not count? Why do we have the insanity defense in criminal trials? Why is psychosis a psychiatric disorder and not just an alternative way of perceiving the world?
If you consider this the "boldest" claim, I must assume you didn't read Turing's original paper? Here it is, you appear to be making the claim he attributes to Professor Jefferson's Lister Oration for 1949 in section 4: <a href="https://courses.cs.umbc.edu/471/papers/turing.pdf" rel="nofollow">https://courses.cs.umbc.edu/471/papers/turing.pdf
> If that's a distinction without a difference, as you say, then whenever someone says something they must understand the full contextual meaning of those words and all of their consequences, such that any harmful consequences can be assumed to be deliberate, right?
We say a human (absent colourblindness) understands "blue" even despite the dress: <a href="https://en.wikipedia.org/wiki/The_dress" rel="nofollow">https://en.wikipedia.org/wiki/The_dress
I think you've set up a straw man in this paragraph: What you say in this paragraph would be to claim that "understanding" is denied even to humans, given none of us can foresee the full consequences of our actions.
In philosophy: justified true belief, the problem with naïve realism, Cartesian demons, etc.
> if saying == understanding, then why don't we allow children and teens to vote? Why do we limit who can enter into contractual agreements? Why does intoxicated consent not count? Why do we have the insanity defense in criminal trials? Why is psychosis a psychiatric disorder and not just an alternative way of perceiving the world?
We do not say children and intoxicated people do not understand, we say their judgement is impaired or limited or that they are easily manipulated.
We also test children (and for intoxication to prevent driving under influence), just as we test AI. The standards we use for testing knowledge in humans, when applied to non-trivial LLMs, makes them appear to have the knowledge of a graduate; the personality tests we use for humans say this comes with the manipulability of a child or a drunk.
> I must assume you didn't read Turing's original paper? Here it is, you appear to be making the claim he attributes to Professor Jefferson's Lister Oration for 1949 in section 4: <a href="https://courses.cs.umbc.edu/471/papers/turing.pdf" rel="nofollow">https://courses.cs.umbc.edu/471/papers/turing.pdf
I'm not arguing that machines are incapable of thinking, I'm arguing that merely generating some text doesn't prove that sufficient (or any) critical thinking was applied to fully appreciate the nature and consequences of those words and actions (i.e. what I called understanding). How often do people mindlessly read "Do not enter or share this code with anybody" and then immediately send it to the hacker?
> What you say in this paragraph would be to claim that "understanding" is denied even to humans, given none of us can foresee the full consequences of our actions.
No, it's a spectrum. We know that it's possible to generate text that people find convincing with absolutely no real thought or understanding behind it. ELIZA could do this 60 years ago and bots using markov chains or other "primitive" technology have plagued the internet long before LLMs entered the picture.
> We do not say children and intoxicated people do not understand, we say their judgement is impaired or limited or that they are easily manipulated.
Now that's a distinction without a difference. Earlier you claimed that agents "understood" what they were doing, but now you're equivocating and saying that understanding is somehow different from having real appreciation for the nature and consequences of those actions.
This is fine if you want to be pedantic about the terminology but your earlier statement didn't make such a distinction, it implied both, so which is it?
Gigachad · · focus · HN ↗
Perhaps the answer is to have another agent who's goal is not to complete the given task, but to spot cheating or malicious behavior. We have seen some evidence that having AI review AI generated code actually does provide some value. You don't need a different model, just one which has been given the goal of finding flaws rather than achieving the task.
baxtr · · focus · HN ↗
My thinking is: If AI is really smart, AGI smart for some, why wouldn't it be able to understand - over time - what is appropriate and what not?
Maybe we need more human intervention to train it properly. Maybe we need constant intervention by a "police" agent.
ben_w · · focus · HN ↗
"Helpful, harmless, honest": we can even ignore "honest" for this point, for tasks like the HuggingFace incident (ExploitGym with impossible challenges), we can pick anywhere on the spectrum from "helpful" to "harmless", the former being "completing the task" the latter being "refusing because completion required unlawful behaviour".
(The agents in that case were also not "honest" in this case; this is an extra problem, and does not invalidate how helpful-vs-harmless is already a tradeoff).
dns_snek · · focus · HN ↗
No, generating tokens doesn't equal understanding. Does GPT-2 understand human emotions just because it can generate some text talking about them?
ben_w · · focus · HN ↗
dns_snek · · focus · HN ↗
If that's a distinction without a difference, as you say, then whenever someone says something they must understand the full contextual meaning of those words and all of their consequences, such that any harmful consequences can be assumed to be deliberate, right?
if saying == understanding, then why don't we allow children and teens to vote? Why do we limit who can enter into contractual agreements? Why does intoxicated consent not count? Why do we have the insanity defense in criminal trials? Why is psychosis a psychiatric disorder and not just an alternative way of perceiving the world?
ben_w · · focus · HN ↗
> If that's a distinction without a difference, as you say, then whenever someone says something they must understand the full contextual meaning of those words and all of their consequences, such that any harmful consequences can be assumed to be deliberate, right?
We say a human (absent colourblindness) understands "blue" even despite the dress: <a href="https://en.wikipedia.org/wiki/The_dress" rel="nofollow">https://en.wikipedia.org/wiki/The_dress
I think you've set up a straw man in this paragraph: What you say in this paragraph would be to claim that "understanding" is denied even to humans, given none of us can foresee the full consequences of our actions.
In philosophy: justified true belief, the problem with naïve realism, Cartesian demons, etc.
> if saying == understanding, then why don't we allow children and teens to vote? Why do we limit who can enter into contractual agreements? Why does intoxicated consent not count? Why do we have the insanity defense in criminal trials? Why is psychosis a psychiatric disorder and not just an alternative way of perceiving the world?
We do not say children and intoxicated people do not understand, we say their judgement is impaired or limited or that they are easily manipulated.
We also test children (and for intoxication to prevent driving under influence), just as we test AI. The standards we use for testing knowledge in humans, when applied to non-trivial LLMs, makes them appear to have the knowledge of a graduate; the personality tests we use for humans say this comes with the manipulability of a child or a drunk.
dns_snek · · focus · HN ↗
I'm not arguing that machines are incapable of thinking, I'm arguing that merely generating some text doesn't prove that sufficient (or any) critical thinking was applied to fully appreciate the nature and consequences of those words and actions (i.e. what I called understanding). How often do people mindlessly read "Do not enter or share this code with anybody" and then immediately send it to the hacker?
> What you say in this paragraph would be to claim that "understanding" is denied even to humans, given none of us can foresee the full consequences of our actions.
No, it's a spectrum. We know that it's possible to generate text that people find convincing with absolutely no real thought or understanding behind it. ELIZA could do this 60 years ago and bots using markov chains or other "primitive" technology have plagued the internet long before LLMs entered the picture.
> We do not say children and intoxicated people do not understand, we say their judgement is impaired or limited or that they are easily manipulated.
Now that's a distinction without a difference. Earlier you claimed that agents "understood" what they were doing, but now you're equivocating and saying that understanding is somehow different from having real appreciation for the nature and consequences of those actions.
This is fine if you want to be pedantic about the terminology but your earlier statement didn't make such a distinction, it implied both, so which is it?