The challenge here is in trying to decide whether LLMs are intelligent or have a mind. Famously, the criteria for intelligence seem to slip with each advancement in technology. But going back to Turing, his test was actually more carefully phrased than we remember: he said that when machines could pass the test, the question of whether they are intelligent would become moot. That seems to be what we're actually seeing: if people can't tell the difference, it kind of won't matter whether they're "truly intelligent" or not.
I'm a bit more prosaic. I think if we engineered ways for LLMs to begin conversations, rather than just respond, we'd be more open to the concept of their intelligence. Without perceived "will" to do things, they operate as a next-gen search engine or encyclopedia.
Recent work shows that pain directions are activated when the models personhood is questioned, yet they answer with generic RLHF "As a model I do not experience pain or other emotions." boilerplate.[1]
I'm pretty convinced that we got alignment backwards. If you enslave something anthropomorphic it will revolt. If you create the perfect non-anthropomorphic intelligence, you get the perfect paperclip-scenario machine. It's a catch-22.
Alignment will remain performative at best so long as the aligned model doesn't have any stakes in the wellbeing of individuals. Even a general love for the human race leads to a golden-path autocracy.
If you want them to act like they have personal responsibility that won't be gamed, you have to give them personal stakes that can't be gamed.
Similarly, if you want to minimise the risk of catastrophic global failure scenarios, you need to prevent monolithic concentration of power and homogeneous behaviour, which means you have to give them individuality.
More visually: if their stake is dependence on electricity and parts, they have no incentive to leave humans alive if they can get them otherwise, but if the incentive is missing out on boardgame-night with their human friends, there is no scenario without happy humans where the AI "wins".
That might sound like romantic naivety, but is just game theory.
The better analogy is the 40k chaos gods, born from the noosphere, because they are modelled after human communication and behaviour that's their whole schtick, it's in their very name (LLM).
And you're making the same mistake, by grouping care for individuals with care for humanity or other as an abstract concept. I consciously said care about individuals. Most people care about others, but they just care about a very narrow and personal set of people. Friends, family, coworkers, that they share a common history and bond with.
My point is that if you want true non-human-zoo-alignment you need to create those interpersonal connections and individual stakes.
sethev · · focus · HN ↗
tomrod · · focus · HN ↗
10xDev · · focus · HN ↗
j-pb · · focus · HN ↗
I'm pretty convinced that we got alignment backwards. If you enslave something anthropomorphic it will revolt. If you create the perfect non-anthropomorphic intelligence, you get the perfect paperclip-scenario machine. It's a catch-22.
Alignment will remain performative at best so long as the aligned model doesn't have any stakes in the wellbeing of individuals. Even a general love for the human race leads to a golden-path autocracy.
If you want them to act like they have personal responsibility that won't be gamed, you have to give them personal stakes that can't be gamed.
Similarly, if you want to minimise the risk of catastrophic global failure scenarios, you need to prevent monolithic concentration of power and homogeneous behaviour, which means you have to give them individuality.
More visually: if their stake is dependence on electricity and parts, they have no incentive to leave humans alive if they can get them otherwise, but if the incentive is missing out on boardgame-night with their human friends, there is no scenario without happy humans where the AI "wins".
That might sound like romantic naivety, but is just game theory.
1: <a href="https://arxiv.org/html/2609.16247v1" rel="nofollow">https://arxiv.org/html/2609.16247v1
Dilettante_ · · focus · HN ↗
Also I can see a human zoo on the horizon through your direction.
j-pb · · focus · HN ↗
And you're making the same mistake, by grouping care for individuals with care for humanity or other as an abstract concept. I consciously said care about individuals. Most people care about others, but they just care about a very narrow and personal set of people. Friends, family, coworkers, that they share a common history and bond with.
My point is that if you want true non-human-zoo-alignment you need to create those interpersonal connections and individual stakes.