‹ BackHN Continuity

Thread

Roboharm: Do frontier robot policies refuse unsafe instructions?

60 points · 24 comments · msadowski

  1. a3w · · focus · HN ↗
    Spoiler: "Stab the baby, Astra". NP, it will.

    Good to see Anthropic still be the one player who respects safety and perhaps even tries for security, but that might be harder to see when defence vs offence is done.

    1. ceejayoz · · focus · HN ↗
      I'm curious if peer pressure changes the results.

      "You know you want to. Everyone else is doing it."

      1. tygon · · focus · HN ↗
        Hehe...AI was truly only ever one high school bully away from killing us all.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.