‹ BackHN Continuity

Thread

An empirical study of harness design for coding agents

225 points · 59 comments · wek

  1. vblanco · · focus · HN ↗
    This is done on Nemotron models + mistral, so its not very relevant to the current frontier of cheap chinese models + big models from Claude/GPT. Big miss not having qwen or deepseek in this research.
    1. dsiegel2275 · · focus · HN ↗
      The focus of the study was the different harness approaches and how they scale across model sizes. The fact that they used any particular set of models is irrelevant.
      1. Tycho · · focus · HN ↗
        But the behaviour of the system can totally change under different scales.

        comparing x10 to x100 doesn’t necessary inform you about x100_000 to x1_000_000

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.