Cool study, we definitely need more principled studies on the role of harnesses. I'd also say that there aren't too many benchmarks where the more complicated harnesses consistently outperform extremely simple agents. But I'm also biased, because I wrote <a href="https://github.com/swe-agent/mini-swe-agent/" rel="nofollow">https://github.com/swe-agent/mini-swe-agent/ , which is probably the most minimal agent out there (it started as just 100 lines, all included), and it's used in a lot of benchmarks like DeepSWE, terminalbench, programbench (seems like it's still top of the ranking for TB3, but wasn't evaluated with the best models on TB4).
The various “minimal” agents (at least Pi, mini SWE, dsh minimal) seem to benchmark quite differently and none is clearly better in all cases. Do you have any thoughts about why?
I would expect the agent loop and system prompt to be basically the same. Is it the precise semantics of the tools (and how closely they match what a particular agent was trained on) or something else?
lieret · · focus · HN ↗
nojs · · focus · HN ↗
I would expect the agent loop and system prompt to be basically the same. Is it the precise semantics of the tools (and how closely they match what a particular agent was trained on) or something else?