You don't need documentation or the 3rd party memory systems. The code IS the documentation.
All this stuff is LLM rube goldberg machines. It just pollutes context.
I barely use AGENTS.md/CLAUDE.md these days. And where they remain, it's super basic high level stuff.
I'm honestly still kicking myself in the ass on many projects where I did something similar to this. I kept tons of markdown docs and decision docs. Now those things are just causing problems because they got stale. Even after having sessions of reconciling documentation, the LLM just gets confused.
Haha, all the software devs who hate writing documentation are naturally finding their preexisting beliefs reinforced when the LLM is able to discern intent without docs. An LLM can be spooky impressive at reading minimized or obfuscated code, for example.
But this article argues that LLMs do better when the context is smaller — when it can understand the totality of the task with as little context as possible. And so having correct API-level docs is greatly advantageous. Anecdotally, this rings true to me — when the local context is good and clear, the LLM writes code matching my intent even when my prompt is sloppy and poorly specified.
Rejoice! The LLM will write the docs for you, relieving you of most of the work.
However without intervention, it will do too much and record absurdly verbose docs (similar to how an LLM will relentlessly refactor your code until you instruct it to move in minimal, incremental changesets). You will still need to edit down what the LLM generates.
100% agree. Most of the things that make development better for humans also make development better for agents… and I think docs are even more important with agents, because of some kind of multiplicative effect. The agents are coding faster, and the benefits of documentation are somewhat more pronounced because of the speed.
The docs end up stale or you waste a ton of time and tokens keeping them up to date. And you can't even trust the LLMs to keep the docs up to date, you will still need to review it yourself keep a bunch of shit out of them.
"docs and maps" are hardly specifying preferences. I DID say I still often keep a high level AGENTS.md or CLAUDE.md. Where I have them, they're super high-level. In places where we have well defined and followed standards that top level markdown file can be as simple as "check out this other repo for patterns" and where I don't, the markdown file is high level overview of what the app is and then a high level overview of architecture or some preferences. Always a shorter/smaller doc.
I don't mind SOME documentation. I'm growing extremely frustrated with the absolute MOUNTAIN of docs, comments, decision documents, etc being created. I'm so fucking tired of the LLM rube goldberg machines being built.
You guys show me a measurable difference on something with docs vs without docs and I'll believe it.
The docs just make you feel good. They're not worth anything. They just create more work if anything because now you don't only have to fight entropy in your codebase but also your docs.
Uh, you aren't providing numbers either, just appending the same assertion to the end of each branch of the discussion. The result is the same as when the LLM does that: geometric increase in what we have to fight through to reason about the problem.
With the LLM, though, I have some influence: I can instruct it to be more succinct. I don't have a specific top-level instruction about that at present in any global prefs file, but I've occasionally asked it to summarize and archive and it's done an okayish job of pruning.
Yes, you're right, I'm not providing numbers. But that's because I'm just using the tool as designed... the vanilla version. I'm using the baseline. I believe anthropic and openai publish plenty for my side of the argument already.
I'm not the one bolting things on claiming it increases productivity. Why would I ADD stuff without any proof or evidence that it works? I'm simply NOT adding these things.
"CLAUDE.md files solve this by giving Claude persistent context about your project. A well-configured CLAUDE.md transforms how Claude works with your specific project. The file serves multiple purposes: providing architectural context, establishing workflows, and connecting Claude to your development tools. Each addition should solve a real problem you have encountered, not theoretical concerns about what Claude might need."
Yeah. Basically IMO/IME the current best practice is to start the LLMs out with minimal/no skills/instructions/docs.
Notice the things they struggle with and the small problems (typically, environment issues IME) they repeatedly encounter and re-solve across multiple sessions. That is what your instructions should cover. When possible, move those instructions into skills, so they get loaded into context selectively instead of on every session. (Example: instructions for running specs, placed into a skill that only gets loaded into context when it’s time to run specs)
Again, this can be automated by the LLMs themselves: both Codex and Claude (and I’m assuming other major harnesses) know how to read their own transcripts and are good at looking for repeated friction and making concrete suggestions to reduce that friction in the future.
It takes a bit of a time investment on the user’s part, and every now and then you probably should throw it all out and start fresh so that the new batch of instructions can be appropriate for the current state of the repo and the capabilities of whatever model(s) you’re using.
Honestly, don't agree. I think even this is a waste of time and context.
The LLMs will always have these suggestions on things that will "reduce friction in the future" but I really think the LLM is a sycophant glazing you. Because even when you add those skills or prompts at some point the LLM ignores it or does whatever it wants to do anyways. Better to leave most of that out and deal with it when it arises.
Even your example, instructions for running specs... none of the current frontier models need this at all.
Nope. Opus and Sol for the most part. Haven't even used Astra or Fable.
At work it's all Anthropic. I spent a while on Sonnet models but have been pretty consistent about using Opus now. For personal projects I bounce between Sol/Terra and Gemini Flash.
Most of the models are good enough now. You don't need all these documents, memory, plugins, skills, etc. MCPs are good for hooking up to external sources of data (at work we have gitlab mcp, datadog mcp, and our knowledge base mcp... which the knowledge base is mostly worthless now because of all this AI generated documentation, but I digress). I DO still use beads on a lot of projects, but not always. My prompts these days are lazy as fuck: "review this repo, check out this other repo with architectural patterns we should follow and libraries/terraform/etc we should use, <basic desc of goal>. What do you think? Let's review everything and discuss before we start building"
Oh no, I might have to stop the llm agent on occasion, or review what they did and tell them to fix a couple things. Way better use of tokens and my time than creating some rube goldberg machine. Way better than keeping a giant decision file that goes stale and causes context corruption because the llm only saw the OLD decision and not the UPDATED decision.
I think they're using Skynet. They seem to be claiming that part of the reason for not giving the LLM directions is because the LLM won't follow them anyway.
That must be a remarkable model. Smart enough to flawlessly discover everything on its own in every session... but also it just straight up just doesn't obey direction.
I'm saying a lot of devs/engineers little rain dances aren't really bringing the rain.
Keep a basic CLAUDE.md/AGENTS.md that keeps high level details. Have a way to reference other projects/apps/codebases to use as reference architecture.
Don't keep tons of markdown files. Don't keep "decision" documents (these are probably the worst). Don't make skills for every little thing (superpowers are dead these days, stop using them). Oh yeah, don't have the LLM generate docs for reference later... if it can generate the docs it doesn't really need them.
I understand what you're railing against in general: engineers who make these big memory/doc systems that quickly go stale and are at best useless and at worst actively harmful. Even more insidiously, these engineers are blind to this fact because maybe some of this stuff was effective 6-12 months ago and they haven't reassessed their workflows. Now that's dogma. We really do agree on that.
However. At least in the projects I've worked on, it's trivial to read the transcripts and notice that without enough direction, there are certain things the agents will burn lots of tokens to rediscover in every session. That's absolutely not dogma.
I think there's a balance between the tokens burned up rediscovering things vs tokens being used keeping docs/decisions/skills/etc.
And I think a lot of places are burning way more on the latter than the former. I think the former is more cost effective due to caching as well. So I HEAVILY lean towards the former especially with how effective I think it is these days.
Don't agree. LLM generated docs are some of the worst because like I said, the LLM never trims, only amends. So we used to do X but now we've had an architectural change or some type of change where we should never do X. Instead of just removing the instructions to do X in the docs, it amends them, "we made the decision to no longer do X because of Y". Now it has X multiple times in context instead of just not having X in context at all.
I hesitate to make blanket statements, since this is quickly evolving, but in general I agree. LLM generated docs will quickly degrade to overly verbose, unreadable crap.
Not just that they get stale, but if you're using the LLM to generate the doc you don't need it.
I don't think I've ever seen an agent review a doc and then NOT also go look at the code. So skip the middleman, just have the agent look at the code.
kaydub · · focus · HN ↗
All this stuff is LLM rube goldberg machines. It just pollutes context.
I barely use AGENTS.md/CLAUDE.md these days. And where they remain, it's super basic high level stuff.
I'm honestly still kicking myself in the ass on many projects where I did something similar to this. I kept tons of markdown docs and decision docs. Now those things are just causing problems because they got stale. Even after having sessions of reconciling documentation, the LLM just gets confused.
rectang · · focus · HN ↗
But this article argues that LLMs do better when the context is smaller — when it can understand the totality of the task with as little context as possible. And so having correct API-level docs is greatly advantageous. Anecdotally, this rings true to me — when the local context is good and clear, the LLM writes code matching my intent even when my prompt is sloppy and poorly specified.
Rejoice! The LLM will write the docs for you, relieving you of most of the work.
However without intervention, it will do too much and record absurdly verbose docs (similar to how an LLM will relentlessly refactor your code until you instruct it to move in minimal, incremental changesets). You will still need to edit down what the LLM generates.
klodolph · · focus · HN ↗
trinsic2 · · focus · HN ↗
kaydub · · focus · HN ↗
The docs end up stale or you waste a ton of time and tokens keeping them up to date. And you can't even trust the LLMs to keep the docs up to date, you will still need to review it yourself keep a bunch of shit out of them.
joquarky · · focus · HN ↗
kaydub · · focus · HN ↗
I don't mind SOME documentation. I'm growing extremely frustrated with the absolute MOUNTAIN of docs, comments, decision documents, etc being created. I'm so fucking tired of the LLM rube goldberg machines being built.
kaydub · · focus · HN ↗
The docs just make you feel good. They're not worth anything. They just create more work if anything because now you don't only have to fight entropy in your codebase but also your docs.
rectang · · focus · HN ↗
With the LLM, though, I have some influence: I can instruct it to be more succinct. I don't have a specific top-level instruction about that at present in any global prefs file, but I've occasionally asked it to summarize and archive and it's done an okayish job of pruning.
kaydub · · focus · HN ↗
I'm not the one bolting things on claiming it increases productivity. Why would I ADD stuff without any proof or evidence that it works? I'm simply NOT adding these things.
frankacter · · focus · HN ↗
While arguing against using the tool as designed.
<a href="https://claude.com/blog/using-claude-md-files" rel="nofollow">https://claude.com/blog/using-claude-md-files
"CLAUDE.md files solve this by giving Claude persistent context about your project. A well-configured CLAUDE.md transforms how Claude works with your specific project. The file serves multiple purposes: providing architectural context, establishing workflows, and connecting Claude to your development tools. Each addition should solve a real problem you have encountered, not theoretical concerns about what Claude might need."
JohnBooty · · focus · HN ↗
Notice the things they struggle with and the small problems (typically, environment issues IME) they repeatedly encounter and re-solve across multiple sessions. That is what your instructions should cover. When possible, move those instructions into skills, so they get loaded into context selectively instead of on every session. (Example: instructions for running specs, placed into a skill that only gets loaded into context when it’s time to run specs)
Again, this can be automated by the LLMs themselves: both Codex and Claude (and I’m assuming other major harnesses) know how to read their own transcripts and are good at looking for repeated friction and making concrete suggestions to reduce that friction in the future.
It takes a bit of a time investment on the user’s part, and every now and then you probably should throw it all out and start fresh so that the new batch of instructions can be appropriate for the current state of the repo and the capabilities of whatever model(s) you’re using.
kaydub · · focus · HN ↗
The LLMs will always have these suggestions on things that will "reduce friction in the future" but I really think the LLM is a sycophant glazing you. Because even when you add those skills or prompts at some point the LLM ignores it or does whatever it wants to do anyways. Better to leave most of that out and deal with it when it arises.
Even your example, instructions for running specs... none of the current frontier models need this at all.
joquarky · · focus · HN ↗
kaydub · · focus · HN ↗
At work it's all Anthropic. I spent a while on Sonnet models but have been pretty consistent about using Opus now. For personal projects I bounce between Sol/Terra and Gemini Flash.
Most of the models are good enough now. You don't need all these documents, memory, plugins, skills, etc. MCPs are good for hooking up to external sources of data (at work we have gitlab mcp, datadog mcp, and our knowledge base mcp... which the knowledge base is mostly worthless now because of all this AI generated documentation, but I digress). I DO still use beads on a lot of projects, but not always. My prompts these days are lazy as fuck: "review this repo, check out this other repo with architectural patterns we should follow and libraries/terraform/etc we should use, <basic desc of goal>. What do you think? Let's review everything and discuss before we start building"
Oh no, I might have to stop the llm agent on occasion, or review what they did and tell them to fix a couple things. Way better use of tokens and my time than creating some rube goldberg machine. Way better than keeping a giant decision file that goes stale and causes context corruption because the llm only saw the OLD decision and not the UPDATED decision.
JohnBooty · · focus · HN ↗
That must be a remarkable model. Smart enough to flawlessly discover everything on its own in every session... but also it just straight up just doesn't obey direction.
Hmmm.
kaydub · · focus · HN ↗
I didn't say give no direction.
I'm saying a lot of devs/engineers little rain dances aren't really bringing the rain.
Keep a basic CLAUDE.md/AGENTS.md that keeps high level details. Have a way to reference other projects/apps/codebases to use as reference architecture.
Don't keep tons of markdown files. Don't keep "decision" documents (these are probably the worst). Don't make skills for every little thing (superpowers are dead these days, stop using them). Oh yeah, don't have the LLM generate docs for reference later... if it can generate the docs it doesn't really need them.
JohnBooty · · focus · HN ↗
kaydub · · focus · HN ↗
You're just stuck on your dogma.
JohnBooty · · focus · HN ↗
However. At least in the projects I've worked on, it's trivial to read the transcripts and notice that without enough direction, there are certain things the agents will burn lots of tokens to rediscover in every session. That's absolutely not dogma.
kaydub · · focus · HN ↗
And I think a lot of places are burning way more on the latter than the former. I think the former is more cost effective due to caching as well. So I HEAVILY lean towards the former especially with how effective I think it is these days.
kaydub · · focus · HN ↗
icedchai · · focus · HN ↗
kaydub · · focus · HN ↗
I don't think I've ever seen an agent review a doc and then NOT also go look at the code. So skip the middleman, just have the agent look at the code.
joquarky · · focus · HN ↗
kaydub · · focus · HN ↗
NBJack · · focus · HN ↗