‹ BackHN Continuity

Thread

HarnessTax: How Much Does the Harness Matter for Coding Agents?

233 points · 99 comments · matt_d

  1. lukax · · focus · HN ↗
    What matters more is that you use the tools that the target model was fine-tuned on.

    E.g. for editing files with Claude models you should use Edit(file_path, old_string, new_string, replace_all) but with GPT models you should use apply_patch_call(patch) (where patch is a custom patch string with custom grammar).

    It appears newer models are better at narive harness tool calls and worse at custom tools that look similar to default tools.

    <a href="https:&#x2F;&#x2F;lucumr.pocoo.org&#x2F;2026&#x2F;7&#x2F;4&#x2F;better-models-worse-tools&#x2F;" rel="nofollow">https:&#x2F;&#x2F;lucumr.pocoo.org&#x2F;2026&#x2F;7&#x2F;4&#x2F;better-models-worse-tools&#x2F;

    1. sn0n · · focus · HN ↗
      Doesn’t that have more to do with the templating of tool-calls and how using them are presented to the models?

      Or is that just why my model likes to break out of the sandbox, going strait to exec shell command and editing files using python on the cli?

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.