I find this - or perhaps the title - a bit surprising.
I've benchmarked GLM 5.3 and DSv4.1-F on my fully-annotated decomp of the Nintendo 3DS's kernel, which I have a good mental understanding of, tasking them to find vulns and other bugs (in Max mode w/ subagents). GLM 5.3 founds almost all the vulns in 30min for $22, while DS only found one vuln for $2 in 40min.
Perhaps DS works better where targets have low-hanging fruits than can be found fast?
If you are talking about publicly known vulns, it's a bit moot since they should be in the training sets. If not, you just burned the vulns to that inference provider's training data (and any intermediary), and future benchmarks will be meaningless.
> If not, you just burned the vulns to that inference provider's training data (and any intermediary), and future benchmarks will be meaningless.
Inference providers can credibly promise to not train on your data if they are in a position to get sued.
TuxSH · · focus · HN ↗
I've benchmarked GLM 5.3 and DSv4.1-F on my fully-annotated decomp of the Nintendo 3DS's kernel, which I have a good mental understanding of, tasking them to find vulns and other bugs (in Max mode w/ subagents). GLM 5.3 founds almost all the vulns in 30min for $22, while DS only found one vuln for $2 in 40min.
Perhaps DS works better where targets have low-hanging fruits than can be found fast?
Aissen · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
[deleted] · · focus · HN ↗
[deleted]
networked · · focus · HN ↗
Inference providers can credibly promise to not train on your data if they are in a position to get sued.