‹ BackHN Continuity

Thread

An empirical study of harness design for coding agents

225 points · 59 comments · wek

  1. rahulmax · · focus · HN ↗
    Quite inline with what I had found with my Claude code sessions over the last year. I wrote about this a few months ago.

    <a href="https:&#x2F;&#x2F;rahulmax.com&#x2F;notes&#x2F;how-i-keep-the-ai-bill-down&#x2F;" rel="nofollow">https:&#x2F;&#x2F;rahulmax.com&#x2F;notes&#x2F;how-i-keep-the-ai-bill-down&#x2F;

    In their case, context management pays off more the tighter your window. Their gap between managing and not managing is 35.7 points of success rate at 32k and 2.7 points at 128k. My version of that was a rule I stick to, as much as I can. I checkpoint a session at about 25-30% of the window, write the state out to a PROGRESS.md and a JSON file of the requirements, and start fresh. This restart costs me 30 seconds, since a bloated session doesn&#x27;t get any cheaper the longer you stay in it.

    Also worth knowing that the models are Nemotron-3 and Mistral-Medium, not the frontier models most people here are paying for.

    1. arcanemachiner · · focus · HN ↗
      I&#x27;ve been pushing the context window well into the 400-600k+ token range lately (mostly Opus 5). I prefer not to since I&#x27;m aware of context collapse, but I&#x27;ve been leaning that way lately.

      Lately, I&#x27;ve been finding that the game of telephone of handoffs causes more mistakes than just letting the session run longer. Not too mention the wasted time waiting for agents to poke around as they bootstrap a new session from the handoff.

      Of course, I still have handoffs for the large-scale plan being accomplished, but I&#x27;ve been having better results letting an agent finishing what it starts. The game of telephone is a painful waste of time when it goes wrong, which is far too often for me.

      1. mmykola87 · · focus · HN ↗

        [dead]

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.