Cool study, we definitely need more principled studies on the role of harnesses. I'd also say that there aren't too many benchmarks where the more complicated harnesses consistently outperform extremely simple agents. But I'm also biased, because I wrote <a href="https://github.com/swe-agent/mini-swe-agent/" rel="nofollow">https://github.com/swe-agent/mini-swe-agent/ , which is probably the most minimal agent out there (it started as just 100 lines, all included), and it's used in a lot of benchmarks like DeepSWE, terminalbench, programbench (seems like it's still top of the ranking for TB3, but wasn't evaluated with the best models on TB4).
Minimal agents also allow room for more focused add-on tools/infra. I'm working on a context management layer[1] and it's very difficult to do well. In our benchmarking, we've observed simple harnesses like Stirrup[2] outperforming more elaborate ones.
lieret · · focus · HN ↗
maxsich · · focus · HN ↗
[1] <a href="https://www.induction.ai/docs/context-management" rel="nofollow">https://www.induction.ai/docs/context-management [2] <a href="https://github.com/ArtificialAnalysis/Stirrup" rel="nofollow">https://github.com/ArtificialAnalysis/Stirrup