‹ BackHN Continuity

Thread

An empirical study of harness design for coding agents

225 points · 59 comments · wek

  1. lieret · · focus · HN ↗
    Cool study, we definitely need more principled studies on the role of harnesses. I&#x27;d also say that there aren&#x27;t too many benchmarks where the more complicated harnesses consistently outperform extremely simple agents. But I&#x27;m also biased, because I wrote <a href="https:&#x2F;&#x2F;github.com&#x2F;swe-agent&#x2F;mini-swe-agent&#x2F;" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;swe-agent&#x2F;mini-swe-agent&#x2F; , which is probably the most minimal agent out there (it started as just 100 lines, all included), and it&#x27;s used in a lot of benchmarks like DeepSWE, terminalbench, programbench (seems like it&#x27;s still top of the ranking for TB3, but wasn&#x27;t evaluated with the best models on TB4).
    1. screamingninja · · focus · HN ↗
      &gt; Minimal: Just some 100 lines of python for the agent class (and a bit more for the environment, model, and run script) — no fancy dependencies!

      It is way more than 100 lines. Why keep advertising something that is no longer the case?

      1. czhu12 · · focus · HN ↗
        I think they are referring to the file <a href="https:&#x2F;&#x2F;github.com&#x2F;SWE-agent&#x2F;mini-swe-agent&#x2F;blob&#x2F;main&#x2F;src&#x2F;minisweagent&#x2F;agents&#x2F;default.py" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;SWE-agent&#x2F;mini-swe-agent&#x2F;blob&#x2F;main&#x2F;src&#x2F;mi...

        190 lines with comments included so close enough

        Which is the actual agent? Everything else is just glue code to connect to different LLM providers, etc

        1. screamingninja · · focus · HN ↗
          190 is not close enough. It is close to double.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.