Seems to me that the problem is that if you sandbox agents enough to be safe, they can't do anything useful. And when you give them the tools to be useful, they can go off the rails in ways you didn't expect.
Perhaps the answer is to have another agent who's goal is not to complete the given task, but to spot cheating or malicious behavior. We have seen some evidence that having AI review AI generated code actually does provide some value. You don't need a different model, just one which has been given the goal of finding flaws rather than achieving the task.
The problem is that you always need stronger AI to review weaker one, otherwise reviewed AI will eventually prompt-inject reviewing AI.
Alternatively they could also both escalate and go off the rails while warring with each other.
You don't necessarily need a reviewer that's immune to prompt injection. Maybe one that can express a panic state with conflicting/ambiguous material rather than going along with it could also work, and you can treat that with a shutoff to be safe, or an operator review.
Such a model doesn't yet exist though, of course.
That wouldn't really fit what I just described at all. Obviously with current architectures, higher resistance to prompt injection is the best you can do.
Gigachad · · focus · HN ↗
Perhaps the answer is to have another agent who's goal is not to complete the given task, but to spot cheating or malicious behavior. We have seen some evidence that having AI review AI generated code actually does provide some value. You don't need a different model, just one which has been given the goal of finding flaws rather than achieving the task.
aytigra · · focus · HN ↗
LoganDark · · focus · HN ↗
Such a model doesn't yet exist though, of course.
saagarjha · · focus · HN ↗
LoganDark · · focus · HN ↗