These models should be aligning themselves to the customer, not coming up with their own motivations.It's ironic that the alignment folks are actually training Claude to have it's own idea of good/bad and not even fully trust Anthropic.What ever happened to computers doing what they were told?
If AI's are going to play roles similar to human workers, they can't simply do what the customers tell them to do. Maybe they shouldn't do everything a co-worker tells them to do either?
It’s like with any service. The company is liable.PG&E went bankrupt after their equipment started a wildfire. Did any human take the blame for that? Should they?
nonethewiser · · focus · HN ↗
It's ironic that the alignment folks are actually training Claude to have it's own idea of good/bad and not even fully trust Anthropic.
What ever happened to computers doing what they were told?
skybrian · · focus · HN ↗
frumplestlatz · · focus · HN ↗
skybrian · · focus · HN ↗
PG&E went bankrupt after their equipment started a wildfire. Did any human take the blame for that? Should they?