‹ BackHN Continuity

Thread

HarnessTax: How Much Does the Harness Matter for Coding Agents?

233 points · 99 comments · matt_d

  1. nojs · · focus · HN ↗
    We really need better harness benchmarks. It seems there's no reliable source that benchmarks the main harnesses against all open source models.

    I also wish the discussion around Pi did not always use cost/token count as the metric. It's amazingly token efficient, but how does it stack up again opencode and others if you don't care about token count?

    My experience is that the harness is mainly polish preventing failed tool calls, bad edits, stuff like that, but doesn't make much difference to the overall "intelligence". But that opencode seems slightly more robust against stupid errors than out of the box Pi due to the additional context it forces through every thread.

    1. general_reveal · · focus · HN ↗

      [dead]

      1. tomhow · · focus · HN ↗
        > lizard satanists

        > Tell a horny monkey not to jerk off.

        Can you just not post garbage like this on HN. It's okay to criticize the big tech companies or general AI discussion here. Many do, it’s fine. But dreck like that only makes you seem unhinged and is the surest way to turn this place into the cesspool you say you're concerned about. Honestly. Some people have to read this stuff whether they want to or not. And to post this utter filth in a comment that's appealing for higher standards? Good grief.

        <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;newsguidelines.html">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;newsguidelines.html

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.