This is done on Nemotron models + mistral, so its not very relevant to the current frontier of cheap chinese models + big models from Claude/GPT. Big miss not having qwen or deepseek in this research.
The focus of the study was the different harness approaches and how they scale across model sizes. The fact that they used any particular set of models is irrelevant.
vblanco · · focus · HN ↗
dsiegel2275 · · focus · HN ↗
siva7 · · focus · HN ↗
So i can skip this study.