Humans missed 1 in 3 threats approving AI agent commands across 40k game runs
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Humans missed 1 in 3 threats approving AI agent commands across 40k game runs
Unofficial Hacker News client; not affiliated with Y Combinator.
continuational · · focus · HN ↗
It's been tried so many times before, and it never worked.
Aurornis · · focus · HN ↗
The common harnesses also have some sandbox functionality, which although imperfect actually does help contain the blast radius for a lot of things.
The common harnesses also support remote development over SSH, which I and many others use to contain development to a virtual machine.
If your complaint is that LLMs can execute tool calls then you’re never going to be happy with any of these solutions and this turns into another generic anti-LLM complaint.
acedTrex · · focus · HN ↗
This is such an unserious approach.
Aurornis · · focus · HN ↗
Like I said above, some people will never be happy with LLMs being allowed to do anything and nothing is going to make them happy about it.
It’s only fair to discuss what the real current status of these systems is. Every time I highlight that things are actually being done, the goalposts move again. There is no possible solution which will satisfy someone who has zero tolerance for letting an LLM execute tool calls because they will always find something.
acedTrex · · focus · HN ↗
Thats fine, theres still a chance it fails.
> There is no possible solution which will satisfy someone who has zero tolerance for letting an LLM execute tool calls because they will always find something.
This is generally correct, security goes completely out of the window with this stuff. It will/currently is a security disaster and theres no actual solution to it.
[deleted] · · focus · HN ↗
[deleted]
wang_li · · focus · HN ↗
Aurornis · · focus · HN ↗
wang_li · · focus · HN ↗
Aurornis · · focus · HN ↗
I think you’re overestimating the revenue generated by this. Having a separate LLM with a cached input prompt check commands is a trivial adder. The only reason it comes up is because they explain to users that it comes out of their plan. So someone on a $20/month plan is going to hit their limits marginally, though mostly negligibly, faster.
If you think they’re sitting in a conference room scheming about making their main models worse on purpose to collect a few extra cents, that’s just baseless conspiracy. They have more to gain or lose based on main model performance.