‹ BackHN Continuity

Thread

Humans missed 1 in 3 threats approving AI agent commands across 40k game runs

213 points · 175 comments · Wirbelwind

  1. continuational · · focus · HN ↗
    It's kinda funny there is still software coming out whose security model is "constantly ask the user for permission, and hope they never make a mistake".

    It's been tried so many times before, and it never worked.

    1. hombre_fatal · · focus · HN ↗
      As opposed to the norm in computing where the average user is expected to just trust rando software, the AI auto-approver that classifies actions the agents wants to take is a huge step up.

      In fact it might actually be the solution that works.

      Imagine if an intelligent agent (in service of the user) had to approve every new outbound connection, system call shape, filesystem command, etc. that arbitrary software wanted to make.

      1. hobofan · · focus · HN ↗
        I do feel like that still needs to add a layer of interactivity to be complete.

        From what I've seen most auto-approvers in coding harnesses either auto-approve or auto-reject, with no middle ground of escalating the decision to the user, and breaking down the pros and cons for the decision.

        1. hombre_fatal · · focus · HN ↗
          Agreed. The experiment is still in its infancy but the direction is great.

          For example, I want to be asked about general shapes/categories of commands as they first appear for a project and then my decision shapes future classification and gets refined and re-scrutinized over time.

          But it gets better every few months. Claude and/or Codex now show a one-line summary for the inline python3 script or grep or pcap command they want to run.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.