These models should be aligning themselves to the customer, not coming up with their own motivations.
They're going to align to the regulator, not the customer. Right now that's Anthropic as they're saying they can self-regulate, and people are willing to let them try, but the landscape could easily move to be regulated by someone else. Moves like speaking to religious leaders is probably a bit of theatre to keep the government at bay.
nonethewiser · · focus · HN ↗
It's ironic that the alignment folks are actually training Claude to have it's own idea of good/bad and not even fully trust Anthropic.
What ever happened to computers doing what they were told?
onion2k · · focus · HN ↗
They're going to align to the regulator, not the customer. Right now that's Anthropic as they're saying they can self-regulate, and people are willing to let them try, but the landscape could easily move to be regulated by someone else. Moves like speaking to religious leaders is probably a bit of theatre to keep the government at bay.
Den_VR · · focus · HN ↗