‹ BackHN Continuity

Thread

An empirical study of harness design for coding agents

225 points · 59 comments · wek

  1. vblanco · · focus · HN ↗
    This is done on Nemotron models + mistral, so its not very relevant to the current frontier of cheap chinese models + big models from Claude/GPT. Big miss not having qwen or deepseek in this research.
    1. Systemerror7A69 · · focus · HN ↗
      I have to say, I am starting to hate this line of reasoning. Yes, LLMs move extremely fast and a lot of improvements are done in a short amount of time.

      And there might be a point to these arguments, vaguely. However:

      There never seems to be - any - kind of counter example or reasoning behind the rationale. You have an in depth and empirical study, done by researchers who, frankly, now their shit (most of the time)

      And on the other hand a random internet comment saying "nope" because...the models aren't the latest.

      If the latest models really would make a difference, you should at least provide some kind of evidence towards that. As it stands though, every time these comments come up this is missing.

      There seems to just be a vaguely defined understanding that "everything changes all the time, and nothing you ever research is transferable to state-of-the-art models"

      Which brings me to my second point about these kinds of arguments:

      LLM models often - aren't - fundamentally different. Yes, they are vastly more capable. And yes, there are emergent properties. But at their core, they function very much similarly. And for quite a while now, there have not been any of these drastic changes we saw when LLMs first become "good enough" for agentic coding.

      I am tired of dismissing empirical evidence and studies every. single. time for reasons without evidence and seemingly a vague sense of "no, but my model is different"

      1. alansaber · · focus · HN ↗
        I sympathise completely but a ~3B parameter model and ~3T parameter model are going to exhibit very different behaviour, one can only infer so much large model behaviour from the former.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.