‹ BackHN Continuity

Thread

Can open-source prompt-injection detectors catch realistic AI agent attacks?

8 points · 4 comments · northbridgedev

  1. perdy · · focus · HN ↗
    Feels like detection at the wrong layer. The injection is text but the damage is a tool call, so the thing worth constraining is which tools the agent can reach and under whose permissions. A detector at 95% still passes one in twenty straight through to an unconstrained tool.
    1. rudratoshs · · focus · HN ↗

      [dead]

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.