‹ BackHN Continuity

Thread

HarnessTax: How Much Does the Harness Matter for Coding Agents?

233 points · 99 comments · matt_d

  1. Supermancho · · focus · HN ↗
    The term "harness" here is being overloaded for the term "agent", which is worrying. Putting that aside, there are many factors that matter. The "harness" context, the execution pattern (parallel vs sequential), the ability to delegate to other models, etc.

    Optimal harnesses use concurrent execution + subagents and are not stuck on one model. Cost and performance are impacted GREATLY by these tactics, regardless of the native agent context (instruction). This kind of single-harness analysis is shallow and misleading, although the finding that "Provider-specific optimization does not guarantee the best pairing" is probably correct, depending on how you measure.

    It is a starting point.

    1. togilvie · · focus · HN ↗
      Starting point is the right framing. Feels like every harness will need tools to identify the best model for the task, best tools, etc. Ultimately, the only metric that matters is cost-per-successful-task, and it's good we're starting to measure these things. But need a lot more discipline on it.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.