‹ BackHN Continuity

Thread

An empirical study of harness design for coding agents

225 points · 59 comments · wek

  1. lieret · · focus · HN ↗
    Cool study, we definitely need more principled studies on the role of harnesses. I&#x27;d also say that there aren&#x27;t too many benchmarks where the more complicated harnesses consistently outperform extremely simple agents. But I&#x27;m also biased, because I wrote <a href="https:&#x2F;&#x2F;github.com&#x2F;swe-agent&#x2F;mini-swe-agent&#x2F;" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;swe-agent&#x2F;mini-swe-agent&#x2F; , which is probably the most minimal agent out there (it started as just 100 lines, all included), and it&#x27;s used in a lot of benchmarks like DeepSWE, terminalbench, programbench (seems like it&#x27;s still top of the ranking for TB3, but wasn&#x27;t evaluated with the best models on TB4).
    1. maxsich · · focus · HN ↗
      Minimal agents also allow room for more focused add-on tools&#x2F;infra. I&#x27;m working on a context management layer[1] and it&#x27;s very difficult to do well. In our benchmarking, we&#x27;ve observed simple harnesses like Stirrup[2] outperforming more elaborate ones.

      [1] <a href="https:&#x2F;&#x2F;www.induction.ai&#x2F;docs&#x2F;context-management" rel="nofollow">https:&#x2F;&#x2F;www.induction.ai&#x2F;docs&#x2F;context-management [2] <a href="https:&#x2F;&#x2F;github.com&#x2F;ArtificialAnalysis&#x2F;Stirrup" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;ArtificialAnalysis&#x2F;Stirrup

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.