‹ BackHN Continuity

Thread

HarnessTax: How Much Does the Harness Matter for Coding Agents?

233 points · 99 comments · matt_d

  1. nojs · · focus · HN ↗
    We really need better harness benchmarks. It seems there's no reliable source that benchmarks the main harnesses against all open source models.

    I also wish the discussion around Pi did not always use cost/token count as the metric. It's amazingly token efficient, but how does it stack up again opencode and others if you don't care about token count?

    My experience is that the harness is mainly polish preventing failed tool calls, bad edits, stuff like that, but doesn't make much difference to the overall "intelligence". But that opencode seems slightly more robust against stupid errors than out of the box Pi due to the additional context it forces through every thread.

    1. sn0n · · focus · HN ↗
      A good benchmark would require a decent number of smaller scoped one off tasks to larger multi step refactors, and also one shot full project of simple to complex varieties. In addition to a series of “conversational” ambiguity filled one-liners.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.