Can open-source prompt-injection detectors catch realistic AI agent attacks?
Thread
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
Can open-source prompt-injection detectors catch realistic AI agent attacks?
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
verdverm · · focus · HN ↗
I have my agents read other instruction files and they don't seem to get affected by the instructions found after a read/bash tool call. Curious if any analysis has been done to see if older prompt injection data sets are even effective anymore.
The whole thing looks heavily agent generated, my trust in them is not very high, how has this been validated or verified by a human?
Should we expect a magic solution in the near future? <a href="https://github.com/rudratoshs/taintgate" rel="nofollow">https://github.com/rudratoshs/taintgate
(side note, it seems my 'no emoji' system prompt line works really well, I forget how obsessed they can be with emojis)
I'm personally setting up to instead use a policy tuned agent on the tool calls themselves (rather than the output), so it never gets run if it has things that it shouldn't be doing. Mainly because they insist on working around instructions that say "don't" or permissions that restrict tools (eg: "git push": "deny" - where they just put the command in a script and run it there, bypassing hard checks)
rudratoshs · · focus · HN ↗
[dead]
williamse · · focus · HN ↗
[dead]
pradeep_kotari · · focus · HN ↗
[dead]
shieldagent · · focus · HN ↗
[dead]
hugopuybareau · · focus · HN ↗
rudratoshs · · focus · HN ↗
[dead]
perdy · · focus · HN ↗
[dead]
rudratoshs · · focus · HN ↗
[dead]
murilomartinspr · · focus · HN ↗
[dead]
textcortex · · focus · HN ↗