This is done on Nemotron models + mistral, so its not very relevant to the current frontier of cheap chinese models + big models from Claude/GPT. Big miss not having qwen or deepseek in this research.
The focus of the study was the different harness approaches and how they scale across model sizes. The fact that they used any particular set of models is irrelevant.
vblanco · · focus · HN ↗
dsiegel2275 · · focus · HN ↗
Tycho · · focus · HN ↗
comparing x10 to x100 doesn’t necessary inform you about x100_000 to x1_000_000