‹ BackHN Continuity

Thread

HarnessTax: How Much Does the Harness Matter for Coding Agents?

233 points · 99 comments · matt_d

  1. nojs · · focus · HN ↗
    We really need better harness benchmarks. It seems there's no reliable source that benchmarks the main harnesses against all open source models.

    I also wish the discussion around Pi did not always use cost/token count as the metric. It's amazingly token efficient, but how does it stack up again opencode and others if you don't care about token count?

    My experience is that the harness is mainly polish preventing failed tool calls, bad edits, stuff like that, but doesn't make much difference to the overall "intelligence". But that opencode seems slightly more robust against stupid errors than out of the box Pi due to the additional context it forces through every thread.

    1. jiaosdjf · · focus · HN ↗
      We really do need better benchmarks and for models too

      - Most people use a harness because of its subscription (most companies pay Anthropic) - All model benchmarks are biased and gamed, harness benchmarks are too few to matter - Everyone is just guessing, acting on sample sizes of 1 and trust me bro vibes

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.