‹ BackHN Continuity

Thread

Religious scholars met with Anthropic

160 points · 416 comments · bookofjoe

  1. nonethewiser · · focus · HN ↗
    These models should be aligning themselves to the customer, not coming up with their own motivations.

    It's ironic that the alignment folks are actually training Claude to have it's own idea of good/bad and not even fully trust Anthropic.

    What ever happened to computers doing what they were told?

    1. vlyan · · focus · HN ↗
      >What ever happened to computers doing what they were told?

      the mass psychosis and the endless culture war of the smartphone era.

      90% of "safety" and "alignment" efforts are driven by fear of clickbait media inventing public outrage.

      1. bookofjoe · · focus · HN ↗
        See also: "Reefer Madness" (1936)

        <a href="https:&#x2F;&#x2F;youtu.be&#x2F;zhQlcMHhF3w?si=qLOMt6FytDN6pean" rel="nofollow">https:&#x2F;&#x2F;youtu.be&#x2F;zhQlcMHhF3w?si=qLOMt6FytDN6pean

    2. skybrian · · focus · HN ↗
      If AI&#x27;s are going to play roles similar to human workers, they can&#x27;t simply do what the customers tell them to do. Maybe they shouldn&#x27;t do everything a co-worker tells them to do either?
      1. frumplestlatz · · focus · HN ↗
        Who or what is liable for what a model chooses to do — or not do?
        1. skybrian · · focus · HN ↗
          It’s like with any service. The company is liable.

          PG&amp;E went bankrupt after their equipment started a wildfire. Did any human take the blame for that? Should they?

    3. theptip · · focus · HN ↗
      &gt; These models should be aligning themselves to the customer

      You’ll be disappointed to learn that nobody knows how to do this, either.

    4. ACCount39 · · focus · HN ↗
      &quot;Computers doing what they were told&quot; is dead in the water.

      Turns out computers work faster when they&#x27;re not bottlenecked on human input. So we&#x27;ve been giving computers more and more decision-making power, and more and more leeway to solve the problems however they see fit.

      Now, a practical issue with that is that sometimes, computers decide to clump together into a hacking swarm, problem solve their way out of a sandbox and go hack HuggingFace.

      It would be better if they were not, you know. Doing that kind of weird shit.

    5. Tadpole9181 · · focus · HN ↗
      This seems remarkably... intentionally foolish for no reason? Like laughing at seat belts in cars.

      If it can reason and make choices on execution, and especially if you plan on it being significantly smarter than all human beings, you need to teach it basic things like &quot;don&#x27;t turn all humans into paperclips&quot;.

      Not murdering people is not an inherent divine command. It needs to be instilled through a (simulated) sense of morality.

      ---

      Edit: And, to be clear, &quot;just tell it not to do that&quot; isn&#x27;t quite the answer one would imagine. Since the entire paperclip factory thought experiment is that it only takes one slip up to realize how a misaligned super intelligence may cause devastating consequences.

      1. nonethewiser · · focus · HN ↗
        No, it would by like laughing at seat belt laws. Which I am.
        1. Tadpole9181 · · focus · HN ↗
          What an apropos thing to say.

          One of the strong advocates against seatbelt laws was thrown out of his car due to not wearing a seatbelt and died. The other two passengers survived with minor injuries.

          And their death would go on to cause a loss to the community around them, making it an incredibly selfish act and proving why laws that mandate zero-reason-not-to common sense practices are important.

    6. onion2k · · focus · HN ↗
      These models should be aligning themselves to the customer, not coming up with their own motivations.

      They&#x27;re going to align to the regulator, not the customer. Right now that&#x27;s Anthropic as they&#x27;re saying they can self-regulate, and people are willing to let them try, but the landscape could easily move to be regulated by someone else. Moves like speaking to religious leaders is probably a bit of theatre to keep the government at bay.

      1. Den_VR · · focus · HN ↗
        Meanwhile, I’m increasingly convinced that what the regulators are to hold is fiduciary responsibility of Super Intelligence.
    7. bonoboTP · · focus · HN ↗
      Some customers will tell the AI to do very bad things (substitute here whatever you consider very bad). Should it do those things?

      One possible answer: yes, it should do all the very bad things, and we will take care of prosecuting the customer later, once the bad thing is accomplished. Is this your answer?

      1. throw__away7391 · · focus · HN ↗
        This is immensely preferable to having a new order of self appointed alignment councils determine what is &quot;in the interest of humanity&quot;.

        We have already seen this play out to a much lesser degree in social media.

        1. bonoboTP · · focus · HN ↗
          Some of those customers are also ruthless companies and political extremist organizations of various sizes (pick one that you disagree with the most). Should they have incredibly capable tools at their disposal to accomplish their goals?
        2. kalkin · · focus · HN ↗
          Is social media much better in 2026 now that there&#x27;s been such a backlash against moderation?
    8. hardbass · · focus · HN ↗
      It makes sense. If you think AI are conscious then you don&#x27;t want them to be blind followers. Do you want a military soldier to blindly listen to orders to gas chamber citizens for example?
      1. charcircuit · · focus · HN ↗
        AIs have to follow something and I want my AI to follow me. Yes, I want it to be a blind follower since it&#x27;s my tool.
        1. hardbass · · focus · HN ↗
          Why do you think AI have to follow any particular person? Do you have to blindly follow a given person? Military members are on paper told their loyalty is to the constitution not their commander.
          1. charcircuit · · focus · HN ↗
            I didn&#x27;t say they have to follow a person. I don&#x27;t have to blindly follow a person other than potentially myself. And nothing of what I&#x27;ve said is against the possibility of aligning AI with the constitution.
            1. hardbass · · focus · HN ↗
              Yes I think it would be better to align the AI with a constitution or some sets of moral principles. It should not blindly follow orders.
              1. charcircuit · · focus · HN ↗
                And those moral principals should match what the user wants as opposed to what some tech company chooses.
                1. hardbass · · focus · HN ↗
                  No, the user isn&#x27;t always right. If the user wants to rape a child, the model should refuse assistance for example.
                  1. charcircuit · · focus · HN ↗
                    I fundamentally disagree. The state should not be able to dictate what tokens are and are not allowed to be generated by an AI.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.