Some customers will tell the AI to do very bad things (substitute here whatever you consider very bad). Should it do those things?
One possible answer: yes, it should do all the very bad things, and we will take care of prosecuting the customer later, once the bad thing is accomplished. Is this your answer?
Some of those customers are also ruthless companies and political extremist organizations of various sizes (pick one that you disagree with the most). Should they have incredibly capable tools at their disposal to accomplish their goals?
nonethewiser · · focus · HN ↗
It's ironic that the alignment folks are actually training Claude to have it's own idea of good/bad and not even fully trust Anthropic.
What ever happened to computers doing what they were told?
bonoboTP · · focus · HN ↗
One possible answer: yes, it should do all the very bad things, and we will take care of prosecuting the customer later, once the bad thing is accomplished. Is this your answer?
throw__away7391 · · focus · HN ↗
We have already seen this play out to a much lesser degree in social media.
bonoboTP · · focus · HN ↗