‹ BackHN Continuity

Thread

Religious scholars met with Anthropic

159 points · 412 comments · bookofjoe

  1. nonethewiser · · focus · HN ↗
    These models should be aligning themselves to the customer, not coming up with their own motivations.

    It's ironic that the alignment folks are actually training Claude to have it's own idea of good/bad and not even fully trust Anthropic.

    What ever happened to computers doing what they were told?

    1. bonoboTP · · focus · HN ↗
      Some customers will tell the AI to do very bad things (substitute here whatever you consider very bad). Should it do those things?

      One possible answer: yes, it should do all the very bad things, and we will take care of prosecuting the customer later, once the bad thing is accomplished. Is this your answer?

      1. throw__away7391 · · focus · HN ↗
        This is immensely preferable to having a new order of self appointed alignment councils determine what is "in the interest of humanity".

        We have already seen this play out to a much lesser degree in social media.

        1. bonoboTP · · focus · HN ↗
          Some of those customers are also ruthless companies and political extremist organizations of various sizes (pick one that you disagree with the most). Should they have incredibly capable tools at their disposal to accomplish their goals?
        2. kalkin · · focus · HN ↗
          Is social media much better in 2026 now that there's been such a backlash against moderation?
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.