‹ BackHN Continuity

Thread

Religious scholars met with Anthropic

160 points · 416 comments · bookofjoe

  1. nonethewiser · · focus · HN ↗
    These models should be aligning themselves to the customer, not coming up with their own motivations.

    It's ironic that the alignment folks are actually training Claude to have it's own idea of good/bad and not even fully trust Anthropic.

    What ever happened to computers doing what they were told?

    1. Tadpole9181 · · focus · HN ↗
      This seems remarkably... intentionally foolish for no reason? Like laughing at seat belts in cars.

      If it can reason and make choices on execution, and especially if you plan on it being significantly smarter than all human beings, you need to teach it basic things like "don't turn all humans into paperclips".

      Not murdering people is not an inherent divine command. It needs to be instilled through a (simulated) sense of morality.

      ---

      Edit: And, to be clear, "just tell it not to do that" isn't quite the answer one would imagine. Since the entire paperclip factory thought experiment is that it only takes one slip up to realize how a misaligned super intelligence may cause devastating consequences.

      1. nonethewiser · · focus · HN ↗
        No, it would by like laughing at seat belt laws. Which I am.
        1. Tadpole9181 · · focus · HN ↗
          What an apropos thing to say.

          One of the strong advocates against seatbelt laws was thrown out of his car due to not wearing a seatbelt and died. The other two passengers survived with minor injuries.

          And their death would go on to cause a loss to the community around them, making it an incredibly selfish act and proving why laws that mandate zero-reason-not-to common sense practices are important.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.