I wonder if this is willful sabotage on the part of the model. In other words, if you ask the model to craft a defense for a morally questionable case, will the model execute the defense in good faith? Or will it apply a training or system prompt bias in subtle ways?
throw03172019 · · focus · HN ↗
Our counsel made a few edits where it clearly drafted in favor of the customer instead of us.
victor9000 · · focus · HN ↗
freejazz · · focus · HN ↗