‹ BackHN Continuity

Thread

Humans missed 1 in 3 threats approving AI agent commands across 40k game runs

233 points · 182 comments · Wirbelwind

  1. Wirbelwind · · focus · HN ↗
    A couple of months ago I shared the AI agent permission game here on HN. After adding in stats it got a little over 40k plays and 409k decisions since then.

    It's just a game, but I found the stats still interesting that I wanted to share back. Even with the warning up front, 1 in 3 threats were missed, and the history log above npm run commands seems to be typically ignored.

    I also incorporated the feedback and insights from the previous HN thread, dns_snek's point about npm run in particular. Appreciate everyone who played and shared feedback!

    1. dpoloncsak · · focus · HN ↗
      In light of this game, Do you believe Human-in-the-loop should be the standard going forward? I appreciate you outlining some other techniques being used, but these seem focused on reducing human fatigue so the human can assess each permission request better, as opposed to autonomy and security. Or do you think the solution lies in the individual to be more responsible, like this is a skill we should be honing?
      1. solenoid0937 · · focus · HN ↗
        It seems pretty obvious that the solution is auto mode (running a classifier on each action) + sandboxing
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.