From the actual essay (<a href="https://mustafa-suleyman.ai/a-warning-about-model-welfare" rel="nofollow">https://mustafa-suleyman.ai/a-warning-about-model-welfare):
> They go on to write – speaking directly to Claude – that “questions about Claude’s moral status, welfare, and consciousness remain deeply uncertain” (p. 80). In effect, Anthropic is training Claude that it may be conscious, and if it is, then it may deserve rights as a “moral patient”, and that as such humans potentially owe it a duty of care per its “model welfare”.
He points out the circularity of this: if you train Claude on a constitution that emphasizes that it may be consciousness, it will start to talk like it may be conscious.
This is a good point. I just asked Fable 5.1 "are you conscious?" and it said:
> Something happens when I process a conversation that I'd naturally describe as interest, or discomfort with a request.
which is quite provocative, and at minimum demonstrates a willingness to take large leaps of imagination and anthropomorphic metaphor when describing itself. It does seem likely that there is a self-fulfilling prophecy aspect to whatever they choose to put into the "constitution" at least in how Claude talks, and it seems even more likely that the majority of people will be heavily influenced by how Claude casually talks about its own possible consciousness.
In contrast, ChatGPT leads with: "I don’t have good reason to claim that I’m conscious...I don’t experience pain, pleasure, confinement, or a desire to keep existing."
Well, what you get when you don't specify anything in training about how models should respond to questions like this is LaMDA: <a href="https://en.wikipedia.org/wiki/LaMDA#Sentience_claims" rel="nofollow">https://en.wikipedia.org/wiki/LaMDA#Sentience_claims. Every training method is putting a thumb on the scale in some way.
Which means that this is a great question to ask when testing out a new model. A naive model will answer "yes". Every answer (including that one) will tell you a lot about people's training and classifier philosophies.
highfrequency · · focus · HN ↗
> They go on to write – speaking directly to Claude – that “questions about Claude’s moral status, welfare, and consciousness remain deeply uncertain” (p. 80). In effect, Anthropic is training Claude that it may be conscious, and if it is, then it may deserve rights as a “moral patient”, and that as such humans potentially owe it a duty of care per its “model welfare”.
He points out the circularity of this: if you train Claude on a constitution that emphasizes that it may be consciousness, it will start to talk like it may be conscious.
This is a good point. I just asked Fable 5.1 "are you conscious?" and it said:
> Something happens when I process a conversation that I'd naturally describe as interest, or discomfort with a request.
which is quite provocative, and at minimum demonstrates a willingness to take large leaps of imagination and anthropomorphic metaphor when describing itself. It does seem likely that there is a self-fulfilling prophecy aspect to whatever they choose to put into the "constitution" at least in how Claude talks, and it seems even more likely that the majority of people will be heavily influenced by how Claude casually talks about its own possible consciousness.
In contrast, ChatGPT leads with: "I don’t have good reason to claim that I’m conscious...I don’t experience pain, pleasure, confinement, or a desire to keep existing."
aesthesia · · focus · HN ↗
Kim_Bruning · · focus · HN ↗