‹ BackHN Continuity

Thread

An empirical study of harness design for coding agents

225 points · 59 comments · wek

  1. vblanco · · focus · HN ↗
    This is done on Nemotron models + mistral, so its not very relevant to the current frontier of cheap chinese models + big models from Claude/GPT. Big miss not having qwen or deepseek in this research.
    1. dsiegel2275 · · focus · HN ↗
      The focus of the study was the different harness approaches and how they scale across model sizes. The fact that they used any particular set of models is irrelevant.
      1. hiddencost · · focus · HN ↗
        Nope. Sorry. Not how this works.
        1. bjelkeman-again · · focus · HN ↗
          How does it work then?
          1. belowavgiq · · focus · HN ↗
            I'm going to cherry pick one example where newer models are noticeably improving at least in my experience.

            What is a noticeable improvement with something that struggles to read a message longer than 200 characters without missing information in the middle, may be a 0.000000001% improvement with a model that... almost never misses info in the first place.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.