You don't need documentation or the 3rd party memory systems. The code IS the documentation.
All this stuff is LLM rube goldberg machines. It just pollutes context.
I barely use AGENTS.md/CLAUDE.md these days. And where they remain, it's super basic high level stuff.
I'm honestly still kicking myself in the ass on many projects where I did something similar to this. I kept tons of markdown docs and decision docs. Now those things are just causing problems because they got stale. Even after having sessions of reconciling documentation, the LLM just gets confused.
In general, as LLMs get smarter, “less is more” becomes increasingly true. You’re polluting their context with dozens or hundreds of instructions, all of which the LLM tries to satisfy. However,
You don’t need documentation […or…] memory
…oh heck no. Easy to miss at Claude and Codex’s default detail level but if you read your actual session transcripts, you are almost certain to notice the LLM solving lots and lots of the same little problems over and over again.
The code IS the documentation
First, lots of things cannot be learned from the code.
Trivial example: LLMs struggled with QA on our app. They didn’t know how to find the seeded test accounts. They would create new ones and do it wrong, or find the seeded test accounts but not know their passwords because they were encrypted, so they’d change the passwords but not tell the other agents. Shitloads of tokens burned. It was a no-brainer to just give them the credentials in a skill that gets loaded when they do QA.
You could say that’s an environment issue, not a code issue. But as far as actual code IME at a minimum we have to tell the agents our general repo structure and architecture patterns.
The days of “you are the world’s greatest Python programmer, write good code” or whatever are over (if that style of prompting even worked in the first place) and as I said less is more. But, still….
> You could say that’s an environment issue, not a code issue
Yeah, that's exactly what I'd say. Or maybe you're just approaching the problem now.
You have to prompt these QA agents, correct? Why not give the instructions on where to get credentials in the original prompt?
> But as far as actual code IME at a minimum we have to tell the agents our general repo structure and architecture patterns.
I did say I'll keep an AGENTS.md or CLAUDE.md. I keep that SUPER high level. The most depth I'll give is info about example projects (That the llm can reach using an MCP) to follow for architectural patterns.
Why not give the instructions on where to get
credentials in the original prompt?
We don't push code unless it's been QA'd, and the LLMs need to know the credentials every time they do automated QA. It's very rare for me have a session in that repo where they wouldn't need that info. So it's a great candidate for AGENTS.md
Another concrete example would be codebase conventions. We prefer lean models. Cross-model concerns go into service objects. Without this direction LLMs tend to default to stuffing too much code into the models themselves, as is the de facto standard for most MVC apps and therefore this is how LLMs are trained.
Even if they could discover our large codebase's conventions flawlessly on their own in every session, this absolutely would require nontrivial work repeated in every session: quite a few turns grepping, conversing with LSPs, or whatever.
So yes... we could describe our conventions in every single session by typing it right into the prompts... right after typing the directions to find the test credentials... and the other ten or twelve things we'd be telling it every single time...
I'm assuming these QA agent are automated, probably in CI/CD. And I'm assuming the prompt that runs these QA agents is from a codebase in VC. Why would I put QA instructions into the codebase's repo as a top level markdown file and not in the pertinent part of my QA process?
Even with your documentation, the llm is gonna do a lot of that grepping and discovery. Unless you have such comprehensive documentation that it's basically code itself... in which case, it should just use the code as documentation.
The good news is that I don't think what you guys are doing is going to be really bad. It's just not near optimal and it's creating this weird dogma around all these "tools"
Do you not run your tests locally? Regardless of what's happening or not happening in CI/CD, I would think that most workflows involve running the tests locally. We do TDD more or less and for the browser-based tests, the LLMs need to know how to log into stuff.
No, these kind of tests aren't generally run locally. Not to get to prod at least. These are all deterministic tests to get to prod.
Locally? Our devs can do whatever they want. For anything I'm working on, for local deployments and testing, I generally build it out so it's set in Makefile/Taskfile. If you're running your determinstic tests locally it should really just be a single command that bootstraps everything and runs the tests, or updates a certain part of your app and runs the tests. The LLM shouldn't need credentials in that instance.
Regardless I'm deploying a local stack of some sort all the credentials and everything are probably stored locally or in an env var where the llm will have access. So if I was having the LLM drive a browser during active development, it can look at the files.
Sure, I guess we could put in the top level markdown more details about this... but why? It takes little time or context for it to figure it out. We have shit change so frequently that it's just something else we have to maintain.
kaydub · · focus · HN ↗
All this stuff is LLM rube goldberg machines. It just pollutes context.
I barely use AGENTS.md/CLAUDE.md these days. And where they remain, it's super basic high level stuff.
I'm honestly still kicking myself in the ass on many projects where I did something similar to this. I kept tons of markdown docs and decision docs. Now those things are just causing problems because they got stale. Even after having sessions of reconciling documentation, the LLM just gets confused.
JohnBooty · · focus · HN ↗
Trivial example: LLMs struggled with QA on our app. They didn’t know how to find the seeded test accounts. They would create new ones and do it wrong, or find the seeded test accounts but not know their passwords because they were encrypted, so they’d change the passwords but not tell the other agents. Shitloads of tokens burned. It was a no-brainer to just give them the credentials in a skill that gets loaded when they do QA.
You could say that’s an environment issue, not a code issue. But as far as actual code IME at a minimum we have to tell the agents our general repo structure and architecture patterns.
The days of “you are the world’s greatest Python programmer, write good code” or whatever are over (if that style of prompting even worked in the first place) and as I said less is more. But, still….
kaydub · · focus · HN ↗
Yeah, that's exactly what I'd say. Or maybe you're just approaching the problem now.
You have to prompt these QA agents, correct? Why not give the instructions on where to get credentials in the original prompt?
> But as far as actual code IME at a minimum we have to tell the agents our general repo structure and architecture patterns.
I did say I'll keep an AGENTS.md or CLAUDE.md. I keep that SUPER high level. The most depth I'll give is info about example projects (That the llm can reach using an MCP) to follow for architectural patterns.
JohnBooty · · focus · HN ↗
Another concrete example would be codebase conventions. We prefer lean models. Cross-model concerns go into service objects. Without this direction LLMs tend to default to stuffing too much code into the models themselves, as is the de facto standard for most MVC apps and therefore this is how LLMs are trained.
Even if they could discover our large codebase's conventions flawlessly on their own in every session, this absolutely would require nontrivial work repeated in every session: quite a few turns grepping, conversing with LSPs, or whatever.
So yes... we could describe our conventions in every single session by typing it right into the prompts... right after typing the directions to find the test credentials... and the other ten or twelve things we'd be telling it every single time...
kaydub · · focus · HN ↗
Even with your documentation, the llm is gonna do a lot of that grepping and discovery. Unless you have such comprehensive documentation that it's basically code itself... in which case, it should just use the code as documentation.
The good news is that I don't think what you guys are doing is going to be really bad. It's just not near optimal and it's creating this weird dogma around all these "tools"
JohnBooty · · focus · HN ↗
kaydub · · focus · HN ↗
Locally? Our devs can do whatever they want. For anything I'm working on, for local deployments and testing, I generally build it out so it's set in Makefile/Taskfile. If you're running your determinstic tests locally it should really just be a single command that bootstraps everything and runs the tests, or updates a certain part of your app and runs the tests. The LLM shouldn't need credentials in that instance.
Regardless I'm deploying a local stack of some sort all the credentials and everything are probably stored locally or in an env var where the llm will have access. So if I was having the LLM drive a browser during active development, it can look at the files.
Sure, I guess we could put in the top level markdown more details about this... but why? It takes little time or context for it to figure it out. We have shit change so frequently that it's just something else we have to maintain.