‹ BackHN Continuity

Thread

Humans missed 1 in 3 threats approving AI agent commands across 40k game runs

208 points · 169 comments · Wirbelwind

  1. xlii · · focus · HN ↗
    I implemented few agent harnesses (and rik! advertising time: <a href="https:&#x2F;&#x2F;rik.axk.sh" rel="nofollow">https:&#x2F;&#x2F;rik.axk.sh), and once doing that I noticed one thing:

    Context-less self-approval is working well. The failure mode is usually false positives (i.e. safe commands being rejected), not the other way around, with root cause of requesting agent underspecifying context (e.g. not mentioning in the request that it&#x27;s made on behalf of user etc.)

    Thus, I&#x27;m running self-approval YOLO modes on state-of-the-art models for quite some time and it didn&#x27;t bit me. It might, but hey, we&#x27;re long gone from the age of predictable software development.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.