You don't need documentation or the 3rd party memory systems. The code IS the documentation.
All this stuff is LLM rube goldberg machines. It just pollutes context.
I barely use AGENTS.md/CLAUDE.md these days. And where they remain, it's super basic high level stuff.
I'm honestly still kicking myself in the ass on many projects where I did something similar to this. I kept tons of markdown docs and decision docs. Now those things are just causing problems because they got stale. Even after having sessions of reconciling documentation, the LLM just gets confused.
Haha, all the software devs who hate writing documentation are naturally finding their preexisting beliefs reinforced when the LLM is able to discern intent without docs. An LLM can be spooky impressive at reading minimized or obfuscated code, for example.
But this article argues that LLMs do better when the context is smaller — when it can understand the totality of the task with as little context as possible. And so having correct API-level docs is greatly advantageous. Anecdotally, this rings true to me — when the local context is good and clear, the LLM writes code matching my intent even when my prompt is sloppy and poorly specified.
Rejoice! The LLM will write the docs for you, relieving you of most of the work.
However without intervention, it will do too much and record absurdly verbose docs (similar to how an LLM will relentlessly refactor your code until you instruct it to move in minimal, incremental changesets). You will still need to edit down what the LLM generates.
Yeah. Basically IMO/IME the current best practice is to start the LLMs out with minimal/no skills/instructions/docs.
Notice the things they struggle with and the small problems (typically, environment issues IME) they repeatedly encounter and re-solve across multiple sessions. That is what your instructions should cover. When possible, move those instructions into skills, so they get loaded into context selectively instead of on every session. (Example: instructions for running specs, placed into a skill that only gets loaded into context when it’s time to run specs)
Again, this can be automated by the LLMs themselves: both Codex and Claude (and I’m assuming other major harnesses) know how to read their own transcripts and are good at looking for repeated friction and making concrete suggestions to reduce that friction in the future.
It takes a bit of a time investment on the user’s part, and every now and then you probably should throw it all out and start fresh so that the new batch of instructions can be appropriate for the current state of the repo and the capabilities of whatever model(s) you’re using.
Honestly, don't agree. I think even this is a waste of time and context.
The LLMs will always have these suggestions on things that will "reduce friction in the future" but I really think the LLM is a sycophant glazing you. Because even when you add those skills or prompts at some point the LLM ignores it or does whatever it wants to do anyways. Better to leave most of that out and deal with it when it arises.
Even your example, instructions for running specs... none of the current frontier models need this at all.
Nope. Opus and Sol for the most part. Haven't even used Astra or Fable.
At work it's all Anthropic. I spent a while on Sonnet models but have been pretty consistent about using Opus now. For personal projects I bounce between Sol/Terra and Gemini Flash.
Most of the models are good enough now. You don't need all these documents, memory, plugins, skills, etc. MCPs are good for hooking up to external sources of data (at work we have gitlab mcp, datadog mcp, and our knowledge base mcp... which the knowledge base is mostly worthless now because of all this AI generated documentation, but I digress). I DO still use beads on a lot of projects, but not always. My prompts these days are lazy as fuck: "review this repo, check out this other repo with architectural patterns we should follow and libraries/terraform/etc we should use, <basic desc of goal>. What do you think? Let's review everything and discuss before we start building"
Oh no, I might have to stop the llm agent on occasion, or review what they did and tell them to fix a couple things. Way better use of tokens and my time than creating some rube goldberg machine. Way better than keeping a giant decision file that goes stale and causes context corruption because the llm only saw the OLD decision and not the UPDATED decision.
I think they're using Skynet. They seem to be claiming that part of the reason for not giving the LLM directions is because the LLM won't follow them anyway.
That must be a remarkable model. Smart enough to flawlessly discover everything on its own in every session... but also it just straight up just doesn't obey direction.
I'm saying a lot of devs/engineers little rain dances aren't really bringing the rain.
Keep a basic CLAUDE.md/AGENTS.md that keeps high level details. Have a way to reference other projects/apps/codebases to use as reference architecture.
Don't keep tons of markdown files. Don't keep "decision" documents (these are probably the worst). Don't make skills for every little thing (superpowers are dead these days, stop using them). Oh yeah, don't have the LLM generate docs for reference later... if it can generate the docs it doesn't really need them.
kaydub · · focus · HN ↗
All this stuff is LLM rube goldberg machines. It just pollutes context.
I barely use AGENTS.md/CLAUDE.md these days. And where they remain, it's super basic high level stuff.
I'm honestly still kicking myself in the ass on many projects where I did something similar to this. I kept tons of markdown docs and decision docs. Now those things are just causing problems because they got stale. Even after having sessions of reconciling documentation, the LLM just gets confused.
rectang · · focus · HN ↗
But this article argues that LLMs do better when the context is smaller — when it can understand the totality of the task with as little context as possible. And so having correct API-level docs is greatly advantageous. Anecdotally, this rings true to me — when the local context is good and clear, the LLM writes code matching my intent even when my prompt is sloppy and poorly specified.
Rejoice! The LLM will write the docs for you, relieving you of most of the work.
However without intervention, it will do too much and record absurdly verbose docs (similar to how an LLM will relentlessly refactor your code until you instruct it to move in minimal, incremental changesets). You will still need to edit down what the LLM generates.
JohnBooty · · focus · HN ↗
Notice the things they struggle with and the small problems (typically, environment issues IME) they repeatedly encounter and re-solve across multiple sessions. That is what your instructions should cover. When possible, move those instructions into skills, so they get loaded into context selectively instead of on every session. (Example: instructions for running specs, placed into a skill that only gets loaded into context when it’s time to run specs)
Again, this can be automated by the LLMs themselves: both Codex and Claude (and I’m assuming other major harnesses) know how to read their own transcripts and are good at looking for repeated friction and making concrete suggestions to reduce that friction in the future.
It takes a bit of a time investment on the user’s part, and every now and then you probably should throw it all out and start fresh so that the new batch of instructions can be appropriate for the current state of the repo and the capabilities of whatever model(s) you’re using.
kaydub · · focus · HN ↗
The LLMs will always have these suggestions on things that will "reduce friction in the future" but I really think the LLM is a sycophant glazing you. Because even when you add those skills or prompts at some point the LLM ignores it or does whatever it wants to do anyways. Better to leave most of that out and deal with it when it arises.
Even your example, instructions for running specs... none of the current frontier models need this at all.
joquarky · · focus · HN ↗
kaydub · · focus · HN ↗
At work it's all Anthropic. I spent a while on Sonnet models but have been pretty consistent about using Opus now. For personal projects I bounce between Sol/Terra and Gemini Flash.
Most of the models are good enough now. You don't need all these documents, memory, plugins, skills, etc. MCPs are good for hooking up to external sources of data (at work we have gitlab mcp, datadog mcp, and our knowledge base mcp... which the knowledge base is mostly worthless now because of all this AI generated documentation, but I digress). I DO still use beads on a lot of projects, but not always. My prompts these days are lazy as fuck: "review this repo, check out this other repo with architectural patterns we should follow and libraries/terraform/etc we should use, <basic desc of goal>. What do you think? Let's review everything and discuss before we start building"
Oh no, I might have to stop the llm agent on occasion, or review what they did and tell them to fix a couple things. Way better use of tokens and my time than creating some rube goldberg machine. Way better than keeping a giant decision file that goes stale and causes context corruption because the llm only saw the OLD decision and not the UPDATED decision.
JohnBooty · · focus · HN ↗
That must be a remarkable model. Smart enough to flawlessly discover everything on its own in every session... but also it just straight up just doesn't obey direction.
Hmmm.
kaydub · · focus · HN ↗
I didn't say give no direction.
I'm saying a lot of devs/engineers little rain dances aren't really bringing the rain.
Keep a basic CLAUDE.md/AGENTS.md that keeps high level details. Have a way to reference other projects/apps/codebases to use as reference architecture.
Don't keep tons of markdown files. Don't keep "decision" documents (these are probably the worst). Don't make skills for every little thing (superpowers are dead these days, stop using them). Oh yeah, don't have the LLM generate docs for reference later... if it can generate the docs it doesn't really need them.