This is done on Nemotron models + mistral, so its not very relevant to the current frontier of cheap chinese models + big models from Claude/GPT. Big miss not having qwen or deepseek in this research.
The focus of the study was the different harness approaches and how they scale across model sizes. The fact that they used any particular set of models is irrelevant.
Agree. Harnesses are effective because they interact with the underlying model effectively. If the latest models were fundamentally different, excluding them would be a miss. But I don’t think they are, at least not in ways that would affect these observations.
I think the confounding issue is that by now, millions of sessions of Claude Code and Codex are now in the training set for these models. So they have been trained to work the way these harnesses are configured, and at least in the case of Claude Code the harness itself is greatly stripped down because the model has absorbed it.
vblanco · · focus · HN ↗
dsiegel2275 · · focus · HN ↗
shermantanktop · · focus · HN ↗
svachalek · · focus · HN ↗