HarnessTax: How Much Does the Harness Matter for Coding Agents?
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
HarnessTax: How Much Does the Harness Matter for Coding Agents?
Unofficial Hacker News client; not affiliated with Y Combinator.
nojs · · focus · HN ↗
I also wish the discussion around Pi did not always use cost/token count as the metric. It's amazingly token efficient, but how does it stack up again opencode and others if you don't care about token count?
My experience is that the harness is mainly polish preventing failed tool calls, bad edits, stuff like that, but doesn't make much difference to the overall "intelligence". But that opencode seems slightly more robust against stupid errors than out of the box Pi due to the additional context it forces through every thread.
sn0n · · focus · HN ↗
leemysw · · focus · HN ↗
[dead]
calgoo · · focus · HN ↗
The harness becomes more and more important, the smaller the model is as you need to offload context management as well as memory to the harness. The big models basically just need a bash prompt tooling and you let the model manage everything inside its own context.
mobelkh · · focus · HN ↗
tacomagick · · focus · HN ↗
verdverm · · focus · HN ↗
disclaimer, I use opencode and have customized parts of it, and will do more, but it is a solid foundation and comes with more out of the box than pi
pi is too minimal for me, I'd go back to my custom built harness if I wanted to be back at that level
esperent · · focus · HN ↗
Such as?
verdverm · · focus · HN ↗
They also have instructions about how to format certain output, which conflicts with the instructions we have in repo. I only discovered yesterday because we were wondering why the agent kept picking certain tools.
esperent · · focus · HN ↗
verdverm · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
sandeepkd · · focus · HN ↗
At this point when all the models have been trained on all available data with the similar algorithm,
1. either you get more data which is not feasible,
2. or get a better algorithm - a possibility ,
3. or write a more targeted harness.
Harnesses for legal, medicine and all are the ones which are getting focus for this reason. Writing benchmarks for these targeted harnesses would be a catching task
general_reveal · · focus · HN ↗
[dead]
tomhow · · focus · HN ↗
> Tell a horny monkey not to jerk off.
Can you just not post garbage like this on HN. It's okay to criticize the big tech companies or general AI discussion here. Many do, it’s fine. But dreck like that only makes you seem unhinged and is the surest way to turn this place into the cesspool you say you're concerned about. Honestly. Some people have to read this stuff whether they want to or not. And to post this utter filth in a comment that's appealing for higher standards? Good grief.
<a href="https://news.ycombinator.com/newsguidelines.html">https://news.ycombinator.com/newsguidelines.html
ebrahimisoheil · · focus · HN ↗
[dead]
tomrod · · focus · HN ↗
I don't think yet that a general benchmark will capture what is needed, but a gain/loss of function along with improvement / loss along the ability of that function seems to be a good path forward.
artdigital · · focus · HN ↗
Personally jumping around a lot to get a feeling for exactly that, and these days liking the Grok harness out of all of them the most
verdverm · · focus · HN ↗
They used to have one chart with {model X harness} for a subset of combos, looks like that is getting an upgrade
jiaosdjf · · focus · HN ↗
- Most people use a harness because of its subscription (most companies pay Anthropic) - All model benchmarks are biased and gamed, harness benchmarks are too few to matter - Everyone is just guessing, acting on sample sizes of 1 and trust me bro vibes