These models should be aligning themselves to the customer, not coming up with their own motivations.It's ironic that the alignment folks are actually training Claude to have it's own idea of good/bad and not even fully trust Anthropic.What ever happened to computers doing what they were told?
> These models should be aligning themselves to the customerYou’ll be disappointed to learn that nobody knows how to do this, either.
nonethewiser · · focus · HN ↗
It's ironic that the alignment folks are actually training Claude to have it's own idea of good/bad and not even fully trust Anthropic.
What ever happened to computers doing what they were told?
theptip · · focus · HN ↗
You’ll be disappointed to learn that nobody knows how to do this, either.