Even simpler: an agent is a simple while loop with tool calls that prompt an LLM continuously. That’s deterministic, standard software. You literally do not have to process tool calls in a way that will execute whatever the model generated. It’s a choice to process a tool call “run_bash” that provides an escape hatch with full execution permissions.
No? We do that everywhere where software is executed. If you only have the choice between full execution permissions and nothing your service is not ready for anything remotely close to production. And you shouldn’t be active in the software industry IMHO
If you give an agent access to any tools that do useful things, it can exploit vulnerabilities. It does not matter what permissions you think you have given it. You do not have any secure software, and LLMs are already better at finding vulnerabilities than we are. If you don't know that, you shouldn't be active in the software industry IMHO.
Alexadar · · focus · HN ↗
dgellow · · focus · HN ↗
We do not have to do that!
jeremyjh · · focus · HN ↗
dgellow · · focus · HN ↗
jeremyjh · · focus · HN ↗