My thought though has always been that I don't want there to be agent-only designated documentation.
I use mattpocock/skills and that generates ADRs (Architectural Decision Records). That only uses skills, including a setup skill that will write a few pointers in AGENTS.md. I always have a CONTRIBUTING.md to document development flow and a CODING_STANDARDS.md. Between those and the README.md and architecture documentation and commit messages the agents seem to be able to find and use docs and keep them up to date. We are also writing a lot of specs and putting those in Github issues.
I’ve been trying to nudge agents (both Claude and GPT) into a red/green/refactor TDD loop and I’m increasingly unconvinced it’s worthwhile. TDD works for humans because it forces us into a pattern of doing the simplest thing that could work, then thinking about how to refactor that. An LLM can be prompted to do that without the skeleton of a failing test, and I plan to experiment with other methods of pushing them towards the same ends.
The value of TDD is it does ensure that tests are written and that will help with future regressions.
However, I found the same problems with agents that it has with humans (usually with humans it isn't TDD but code coverage requirements). The problem is the tests are now written to satisfy a bureaucracy rather than to properly verify the code. And there are real studies that point out the low value of these types of bureaucratic testing requirements. For example, when models are just required to write a test file, they don't produce a better result: <a href="https://arxiv.org/abs/2602.07900" rel="nofollow">https://arxiv.org/abs/2602.07900
I wrote a /verify skill that focuses on properly verifying code and I am finding it works a lot better. I do need to revise it now- in practice certain parts of the skill are doing all the heavy lifting and others are more dead weight. But it does seem to be properly orienting the agents towards finding defects.
<a href="https://github.com/gregwebs/skills-sdlc/blob/main/skills/verify/SKILL.md" rel="nofollow">https://github.com/gregwebs/skills-sdlc/blob/main/skills/ver...
don't forget: tests also inform the agent on what the code is _supposed_ to do. So with the cycle, you end up with 3 different potentially robust ways to look at and understand how to extend.
Context then is: docs, code and tests. Rather than try to build some omniscient agent, the context is built where the agent needs it.
I agree that ensuring testing is the basis for modifying code in _future_ development (including understanding it). My problem is that we can't produce a study showing that TDD is producing better results _now_.
TDD doesn't help the agent figure out what it is supposed to do now- it is already writing the code and now and it must know what it is supposed to do to write it.
I am trying to use evidence-based approaches. I am only observing better results and tweaking now, but I plan to benchmark when tweaking is done.
gregwebs · · focus · HN ↗
My thought though has always been that I don't want there to be agent-only designated documentation.
I use mattpocock/skills and that generates ADRs (Architectural Decision Records). That only uses skills, including a setup skill that will write a few pointers in AGENTS.md. I always have a CONTRIBUTING.md to document development flow and a CODING_STANDARDS.md. Between those and the README.md and architecture documentation and commit messages the agents seem to be able to find and use docs and keep them up to date. We are also writing a lot of specs and putting those in Github issues.
cyanydeez · · focus · HN ↗
Once qwen3.8-flash-next showed up, it cañ go "forever" with dynamic context pruning.
It's fascinating for local coding.
jon-wood · · focus · HN ↗
gregwebs · · focus · HN ↗
However, I found the same problems with agents that it has with humans (usually with humans it isn't TDD but code coverage requirements). The problem is the tests are now written to satisfy a bureaucracy rather than to properly verify the code. And there are real studies that point out the low value of these types of bureaucratic testing requirements. For example, when models are just required to write a test file, they don't produce a better result: <a href="https://arxiv.org/abs/2602.07900" rel="nofollow">https://arxiv.org/abs/2602.07900
I wrote a /verify skill that focuses on properly verifying code and I am finding it works a lot better. I do need to revise it now- in practice certain parts of the skill are doing all the heavy lifting and others are more dead weight. But it does seem to be properly orienting the agents towards finding defects. <a href="https://github.com/gregwebs/skills-sdlc/blob/main/skills/verify/SKILL.md" rel="nofollow">https://github.com/gregwebs/skills-sdlc/blob/main/skills/ver...
cyanydeez · · focus · HN ↗
Context then is: docs, code and tests. Rather than try to build some omniscient agent, the context is built where the agent needs it.
gregwebs · · focus · HN ↗
TDD doesn't help the agent figure out what it is supposed to do now- it is already writing the code and now and it must know what it is supposed to do to write it.
I am trying to use evidence-based approaches. I am only observing better results and tweaking now, but I plan to benchmark when tweaking is done.
cyanydeez · · focus · HN ↗
[dead]