‹ BackHN Continuity

Thread

HarnessTax: How Much Does the Harness Matter for Coding Agents?

233 points · 99 comments · matt_d

  1. nojs · · focus · HN ↗
    We really need better harness benchmarks. It seems there's no reliable source that benchmarks the main harnesses against all open source models.

    I also wish the discussion around Pi did not always use cost/token count as the metric. It's amazingly token efficient, but how does it stack up again opencode and others if you don't care about token count?

    My experience is that the harness is mainly polish preventing failed tool calls, bad edits, stuff like that, but doesn't make much difference to the overall "intelligence". But that opencode seems slightly more robust against stupid errors than out of the box Pi due to the additional context it forces through every thread.

    1. sandeepkd · · focus · HN ↗
      Harness and benchmark for the harness feels like a chicken and egg problem. The harness is to optimize the interaction results with the models. Any benchmark for harness has to focus on the goals that the harness was trying to optimize for unless we are only focussing on generic harnesses.

      At this point when all the models have been trained on all available data with the similar algorithm,

      1. either you get more data which is not feasible,

      2. or get a better algorithm - a possibility ,

      3. or write a more targeted harness.

      Harnesses for legal, medicine and all are the ones which are getting focus for this reason. Writing benchmarks for these targeted harnesses would be a catching task

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.