‹ BackHN Continuity

Thread

HarnessTax: How Much Does the Harness Matter for Coding Agents?

233 points · 99 comments · matt_d

  1. lukax · · focus · HN ↗
    What matters more is that you use the tools that the target model was fine-tuned on.

    E.g. for editing files with Claude models you should use Edit(file_path, old_string, new_string, replace_all) but with GPT models you should use apply_patch_call(patch) (where patch is a custom patch string with custom grammar).

    It appears newer models are better at narive harness tool calls and worse at custom tools that look similar to default tools.

    <a href="https:&#x2F;&#x2F;lucumr.pocoo.org&#x2F;2026&#x2F;7&#x2F;4&#x2F;better-models-worse-tools&#x2F;" rel="nofollow">https:&#x2F;&#x2F;lucumr.pocoo.org&#x2F;2026&#x2F;7&#x2F;4&#x2F;better-models-worse-tools&#x2F;

    1. imtringued · · focus · HN ↗
      This is correct. People seem to get the wrong idea about why agentic coding is even a thing in 2026. The naive AI techno optimist which has basically displaced the vast majority of opinions on HN, thinks that the models got &quot;smarter&quot; [0]. No, the training distribution shifted towards training on agentic sessions which made certain forms of agentic coding &quot;in-distribution&quot;.

      We are still witnessing the same underlying problems of transformers.

      [0] Think back to all the publicity stunts like the Hugging Face. They are meant to convince you that the agents have somehow progressed past the transformer limitations when those publicity stunts are actually expressions of transformer limitations.

      1. TedDoesntTalk · · focus · HN ↗
        You think the hugging face incident was a stunt? Can you explain?
        1. atwrk · · focus · HN ↗
          OpenAI started fearmongering way back with GPT 2, arguing that model was too dangerous to release freely. That model was barely coherent enough for using it as a twitter bot. Anthropic just hopped onto that later. Conveniently, calling for regulation now would ease the competition from open Chinese models, opening the chance for both companies to eventually reach positive ROI, with consumers paying the price.

          The burden of proof that this isn&#x27;t just a publicity stunt again is squarely on them.

          1. nottorp · · focus · HN ↗

            [dead]

      2. wollowollo · · focus · HN ↗
        There&#x27;s a widespread perception that models are worse for prose and creative writing now. That would track.
      3. slopinthebag · · focus · HN ↗
        yea, plus they have been rlhf&#x27;ed to an inch of their lives as well. hard to tell if frontier models can solve more problems because of that or not.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.