‹ BackHN Continuity

Thread

Humans missed 1 in 3 threats approving AI agent commands across 40k game runs

210 points · 169 comments · Wirbelwind

  1. VladVladikoff · · focus · HN ↗
    I remember when this game was posted here, and there was a lot of discussion at the time that some of the prompts were misleading about whether or not they were risky, some people were debating about how some of the prompts flagged as bad weren’t bad, and others flagged as not bad were. This is a fundamental flaw in the test, that makes the analysis of results meaningless.

    Also the game was on a timer, and maybe there are some very abusive workplaces where you feel that kind of pressure, but I think most of us actually take the time to understand what a being asked before approving it.

    1. thayne · · focus · HN ↗
      Also a lot of them may or may not be safe depending on additional context that you don't have in the test.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.