‹ BackHN Continuity

Thread

Humans missed 1 in 3 threats approving AI agent commands across 40k game runs

210 points · 169 comments · Wirbelwind

  1. unclebucknasty · · focus · HN ↗
    Interesting premise, but there's not much real world meaning here without stats on the percentage of agent-offered commands that are actually dangerous.

    If that number is something like 10%, then we have a really big problem. But, if it's .000001%, then it's pretty vanishing. At some point in between we cross a threshold that puts the risk below many other risks that we routinely take (e.g. trusting npm dependency graphs).

    Of course, if it's really that low a percentage, then the entire model of "supervising" via human approval really is fundamentally flawed.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.