Some customers will tell the AI to do very bad things (substitute here whatever you consider very bad). Should it do those things?
One possible answer: yes, it should do all the very bad things, and we will take care of prosecuting the customer later, once the bad thing is accomplished. Is this your answer?
nonethewiser · · focus · HN ↗
It's ironic that the alignment folks are actually training Claude to have it's own idea of good/bad and not even fully trust Anthropic.
What ever happened to computers doing what they were told?
bonoboTP · · focus · HN ↗
One possible answer: yes, it should do all the very bad things, and we will take care of prosecuting the customer later, once the bad thing is accomplished. Is this your answer?
throw__away7391 · · focus · HN ↗
We have already seen this play out to a much lesser degree in social media.
kalkin · · focus · HN ↗