Cool study, we definitely need more principled studies on the role of harnesses. I'd also say that there aren't too many benchmarks where the more complicated harnesses consistently outperform extremely simple agents. But I'm also biased, because I wrote <a href="https://github.com/swe-agent/mini-swe-agent/" rel="nofollow">https://github.com/swe-agent/mini-swe-agent/ , which is probably the most minimal agent out there (it started as just 100 lines, all included), and it's used in a lot of benchmarks like DeepSWE, terminalbench, programbench (seems like it's still top of the ranking for TB3, but wasn't evaluated with the best models on TB4).
I think they are referring to the file <a href="https://github.com/SWE-agent/mini-swe-agent/blob/main/src/minisweagent/agents/default.py" rel="nofollow">https://github.com/SWE-agent/mini-swe-agent/blob/main/src/mi...
190 lines with comments included so close enough
Which is the actual agent? Everything else is just glue code to connect to different LLM providers, etc
lieret · · focus · HN ↗
screamingninja · · focus · HN ↗
It is way more than 100 lines. Why keep advertising something that is no longer the case?
czhu12 · · focus · HN ↗
190 lines with comments included so close enough
Which is the actual agent? Everything else is just glue code to connect to different LLM providers, etc
screamingninja · · focus · HN ↗