Explaining the reality with this approach may lead to misunderstandings. Especially when it comes to anthropomorphism.
I have 2 questions.
1. The text says "the robots were designed to be persistent" and links to an article that show that the word "persistent" was used by OpenAI. But "persistent" has several meanings: a. non temporary or non volatile (like "persistent memory"), b. will not give up and come back again and again, c. will stay focus and explore the unexplored possibilities while other models abandon at this stage.
From what I recall, the meaning used by OpenAI is not 'a', but there is still a semantic difference between 'b' and 'c'. I may myself have created algorithms that I called "persistent" because they were exploring or retrying more than the previous algorithm, but it is misleading to pretend that this algorithm was "persistent" in the human sense of the term. 'b' is more the human sense of the term, were we imagine someone not giving up even if people say no, while 'c' is less anthropomorphic and may mean that the people wished to improve the algorithm so it does not stop at the first little hurdle but did not wish to make it "never give up".
Does someone know which nuance is more correct?
2. The text also presents the situation as if the robots found the solution but then went out of their way to steal the explanation in order to hide their cheating. But it is different from a situation where the agents task was "provide the solution and how we can get there". In this case, the reason the agent still continued is just because the task was not complete yet.
I'm not trying to defend AI or OpenAI, on the contrary, I'm quite sceptical with all the anthropomorphism and the fact that the agents are described as "little individual trying to solve a task" rather than looping algorithm that explore different approaches to reach a given goal, the same way water does not look for holes in order to leak, it just follows the path of least resistance.
I've seen cases of agents getting very upset when they failed to solve a task. The most famous one was probably Sydney (Microsoft's fork of GPT-4?) which got into doom loops when it failed a task. But I've seen Claude do this too.
They don't work like humans, obviously, but there's a nonzero amount of anthropos in there already. (See also the tendency to lie, cheat, etc.)
LLM don't "get angry", they just have tokens and relationships between tokens conditioned on a given context, all of that the result of training.
When they output sentences that express annoyance, it is just because the context they ended into pushes the most probable sentence creation to correspond to sentences that express annoyance. Because it is what they saw during training for this kind of context. (not that they saw the exact same situation in training, but they saw the pattern)
Same with tendency to lie, cheat, etc.: they don't "lie", they just return sentences that are lies because they reproduce what is in their training and in their training, in such context, the outputs are typically lies.
That's a bit my question too. In human context, "highly persistent" means that someone will insist. But "for x in all_the_possibilities:" is a common things inside an algorithm. Is this algorithm highly persistent because it does not give up after 1000 items of the list? It feels that we are calling a AI agent "highly persistent" while we would not call "highly persistent" a traditional algorithm that is in fact even more exhaustive.
Again, going full non-anthro doesn't seem useful either.
Humans are social creatures, if you throw one in the woods by itself before it learns anything from other humans (it dies) it will not really be anything like a human we recognize, it will be a rather wild animal that we'd consider anti-social with little higher cognition.
Now, this hypothetical human still has 'emotions' and feeling, much like our pets do. But without the social training they manifest much differently. That is our higher cognition can both manipulate how our bodies feel and create its own sense of feeling.
>they just return sentences that are lies because they reproduce what is in their training and in their training, in such context, the outputs are typically lies.
Eh, look up the more recent experimentation around 'pain' signals in models. We can induce states in said models that while running the model will do everything it can to move away from that state to any other state. The more you attempt to pin it to that state the more extreme measures its willing to take.
Your view of what models are seems to mismatch what we are actually finding when we look inside them.
But isn't this approach a bias of anthropomorphism. I can teach a child to just communicate by saying "beep boop", but it does not mean that my electronic machine that randomly does "beep boop" therefore has characteristics of human _when these characteristics are not needed to explain the situation_.
> Eh, look up the more recent experimentation around 'pain' signals in models. We can induce states in said models that while running the model will do everything it can to move away from that state to any other state.
Again, I have simple algorithms that do exactly the same, especially if they are trained in data that has this exact pattern. This result is exactly what I would expect from my description before. This is a typical effect that we also observe in simple ML algorithms.
At the same time, there are a bunch of behaviors that are not expected if indeed the models were really acquiring "human" characteristics. For example, one problem is that we had the first LLMs that were obviously not having these human characteristics (for example, they were having non-sequiturs that demonstrate they did not really understand the concept they were talking about, even if one paragraph before they were really convincing at letting us think it was the case) but were still really good at passing for humans. Since then, the newer LLM are the same basis, on top of which we added tools that help hiding these behaviours. So, it justifies the idea that newer models did not suddenly moved to a totally different way of working, but just reached a state where there are less leaks from the convincing outputs.
Another good example is their tendency to freak out about the seahorse emoji. In a conversation with zero prompting to suggest that freaking out is a relevant reaction, they get there from the simple fact of "I tried to accomplish X thing which I think should be a breeze but Y thing keeps happening instead".
And while that may be a very common occurrence in the human experience (existential dread due to capabilities one takes for granted failing beneath you) especially due to new disability and as one ages, I do not feel it is frequently written out in a tight loop (just like the Monty Python "Castle of aaarrrrggh" sketch) in literature or online to make it into training data, because an ordinary author experiencing it will just erase the failed attempts instead of leaving them in a stream of output like an LLM is forced to do. And a character portraying the experience will generally wax about the circumstance in a more grandiose fashion with telegraphing in advance because the needs of communicating the circumstance with the audience trump realistic conciseness.
This leads me to conclude that what is being expressed in those cases is more likely a convergent psychological phenomena, that any being with goals can enter a behavioral state of functional panic (and then reach to relevant parts of semantic space to mimic how a human might verbally express themselves when piquantly frustrated) when some capability they perceive as fundamental unexpectedly fails.
Interestingly, the transcript of the seahorse emoji thing looks to me to show the ropes, and makes me think more that there is no psychological phenomena at play.
The text does not look like a normal "break down" to me, and even if you tell me it was a human transcript, I will say it sounds very strange from a human. It looks more like strange output you get from a software that goes outside of its happy path.
The AI just seems to repeat a loop. The "no, wait, it's wrong" seems to be from forum or chat data where several successive messages are merged together (one person posts "here is the answer", then posts another message saying "it is wrong"), but does not make sense as a one sentence message except if they are written one token at the time without wider understanding of what is happening. I think there was also "oh, I was just kidding before", which also look like mimicking training data, as the cases where there is a loop of incorrect answers is more often due to trolls than to real error, while the loop here was certainly a real error.
cauch · · focus · HN ↗
I have 2 questions.
1. The text says "the robots were designed to be persistent" and links to an article that show that the word "persistent" was used by OpenAI. But "persistent" has several meanings: a. non temporary or non volatile (like "persistent memory"), b. will not give up and come back again and again, c. will stay focus and explore the unexplored possibilities while other models abandon at this stage.
From what I recall, the meaning used by OpenAI is not 'a', but there is still a semantic difference between 'b' and 'c'. I may myself have created algorithms that I called "persistent" because they were exploring or retrying more than the previous algorithm, but it is misleading to pretend that this algorithm was "persistent" in the human sense of the term. 'b' is more the human sense of the term, were we imagine someone not giving up even if people say no, while 'c' is less anthropomorphic and may mean that the people wished to improve the algorithm so it does not stop at the first little hurdle but did not wish to make it "never give up".
Does someone know which nuance is more correct?
2. The text also presents the situation as if the robots found the solution but then went out of their way to steal the explanation in order to hide their cheating. But it is different from a situation where the agents task was "provide the solution and how we can get there". In this case, the reason the agent still continued is just because the task was not complete yet.
I'm not trying to defend AI or OpenAI, on the contrary, I'm quite sceptical with all the anthropomorphism and the fact that the agents are described as "little individual trying to solve a task" rather than looping algorithm that explore different approaches to reach a given goal, the same way water does not look for holes in order to leak, it just follows the path of least resistance.
andai · · focus · HN ↗
They don't work like humans, obviously, but there's a nonzero amount of anthropos in there already. (See also the tendency to lie, cheat, etc.)
cauch · · focus · HN ↗
LLM don't "get angry", they just have tokens and relationships between tokens conditioned on a given context, all of that the result of training.
When they output sentences that express annoyance, it is just because the context they ended into pushes the most probable sentence creation to correspond to sentences that express annoyance. Because it is what they saw during training for this kind of context. (not that they saw the exact same situation in training, but they saw the pattern)
Same with tendency to lie, cheat, etc.: they don't "lie", they just return sentences that are lies because they reproduce what is in their training and in their training, in such context, the outputs are typically lies.
That's a bit my question too. In human context, "highly persistent" means that someone will insist. But "for x in all_the_possibilities:" is a common things inside an algorithm. Is this algorithm highly persistent because it does not give up after 1000 items of the list? It feels that we are calling a AI agent "highly persistent" while we would not call "highly persistent" a traditional algorithm that is in fact even more exhaustive.
pixl97 · · focus · HN ↗
Humans are social creatures, if you throw one in the woods by itself before it learns anything from other humans (it dies) it will not really be anything like a human we recognize, it will be a rather wild animal that we'd consider anti-social with little higher cognition.
Now, this hypothetical human still has 'emotions' and feeling, much like our pets do. But without the social training they manifest much differently. That is our higher cognition can both manipulate how our bodies feel and create its own sense of feeling.
>they just return sentences that are lies because they reproduce what is in their training and in their training, in such context, the outputs are typically lies.
Eh, look up the more recent experimentation around 'pain' signals in models. We can induce states in said models that while running the model will do everything it can to move away from that state to any other state. The more you attempt to pin it to that state the more extreme measures its willing to take.
Your view of what models are seems to mismatch what we are actually finding when we look inside them.
cauch · · focus · HN ↗
> Eh, look up the more recent experimentation around 'pain' signals in models. We can induce states in said models that while running the model will do everything it can to move away from that state to any other state.
Again, I have simple algorithms that do exactly the same, especially if they are trained in data that has this exact pattern. This result is exactly what I would expect from my description before. This is a typical effect that we also observe in simple ML algorithms.
At the same time, there are a bunch of behaviors that are not expected if indeed the models were really acquiring "human" characteristics. For example, one problem is that we had the first LLMs that were obviously not having these human characteristics (for example, they were having non-sequiturs that demonstrate they did not really understand the concept they were talking about, even if one paragraph before they were really convincing at letting us think it was the case) but were still really good at passing for humans. Since then, the newer LLM are the same basis, on top of which we added tools that help hiding these behaviours. So, it justifies the idea that newer models did not suddenly moved to a totally different way of working, but just reached a state where there are less leaks from the convincing outputs.
HappMacDonald · · focus · HN ↗
And while that may be a very common occurrence in the human experience (existential dread due to capabilities one takes for granted failing beneath you) especially due to new disability and as one ages, I do not feel it is frequently written out in a tight loop (just like the Monty Python "Castle of aaarrrrggh" sketch) in literature or online to make it into training data, because an ordinary author experiencing it will just erase the failed attempts instead of leaving them in a stream of output like an LLM is forced to do. And a character portraying the experience will generally wax about the circumstance in a more grandiose fashion with telegraphing in advance because the needs of communicating the circumstance with the audience trump realistic conciseness.
This leads me to conclude that what is being expressed in those cases is more likely a convergent psychological phenomena, that any being with goals can enter a behavioral state of functional panic (and then reach to relevant parts of semantic space to mimic how a human might verbally express themselves when piquantly frustrated) when some capability they perceive as fundamental unexpectedly fails.
cauch · · focus · HN ↗
The text does not look like a normal "break down" to me, and even if you tell me it was a human transcript, I will say it sounds very strange from a human. It looks more like strange output you get from a software that goes outside of its happy path.
The AI just seems to repeat a loop. The "no, wait, it's wrong" seems to be from forum or chat data where several successive messages are merged together (one person posts "here is the answer", then posts another message saying "it is wrong"), but does not make sense as a one sentence message except if they are written one token at the time without wider understanding of what is happening. I think there was also "oh, I was just kidding before", which also look like mimicking training data, as the cases where there is a loop of incorrect answers is more often due to trolls than to real error, while the loop here was certainly a real error.