‹ BackHN Continuity

Thread

An empirical study of harness design for coding agents

225 points · 59 comments · wek

  1. agentdev001 · · focus · HN ↗
    As far as I can tell, the paper says "bash capable", without ever describing what that means. How would one know whether a given model is "bash capable" or not?

    I would have to imagine, that Luna would very much fall into the camp of "bash capable". At which point- it seems to me that adding any tools beyond just Bash requires some rigorous testing and verification that value is being added.

    1. everforward · · focus · HN ↗
      I think it’s sort of self-defined. If a model is able to use bash well enough to not need specific tools.

      The research seems to agree with you, though. The paper calls out that for “bash capable” models, adding tools to do things bash can already do doesn’t improve performance.

      Vaguely the same result as RAG. Unless you’re in specific domains, you won’t beat handing the agent a shell and grep.

      1. reddit_clone · · focus · HN ↗
        I am still wondering about the effectiveness of using grep/awk (and ad-hoc python scripts) in code bases, as opposed to more sophisticated LSP and the like?
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.