Agents don't need memory, they need documentation
Thread
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
Agents don't need memory, they need documentation
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
michaelastreiko · · focus · HN ↗
[dead]
jdw64 · · focus · HN ↗
However, AI works differently from humans in that much more of its working context has to be made explicit. Because of that, there may be some fundamentally different way for AI to maintain or reconstruct a program’s overall model.
iamwil · · focus · HN ↗
jdw64 · · focus · HN ↗
[dead]
nottorp · · focus · HN ↗
LLMs don't retrain, it's too expensive atm.
alienbaby · · focus · HN ↗
however, I have found keeping a good solid reference to my home infrastructure, services, ci/cd setup, hosts, storage , networking etc.. really works wonders as a set of 'memories' to share across projects that I expect to be tested / deployed / acceptance tested etc.. using the home infra bits and pieces.
kaydub · · focus · HN ↗
The code is the documentation. There's no need for most of this stuff.
stavros · · focus · HN ↗
<a href="https://github.com/skorokithakis/gnosis" rel="nofollow">https://github.com/skorokithakis/gnosis
This just does a pass at the end, rather than consume context all the time. I've been using it for months and it works great.
espeed · · focus · HN ↗
nextaccountic · · focus · HN ↗
Github Copilot had the idea of attaching memory to files, and if the file hash changes the memory is automatically dropped (not sure if they still do it). This means they are overly eager to drop stuff (even if the file change is just cosmetic), but at least they don't accumulate outdated cruft too much. (a memory can still be outdated if it was invalidated by a change in another file though)
espeed · · focus · HN ↗
monneyboi · · focus · HN ↗
You have the whole session history right there. One recall skill and some JSON parsing gets you grep over perfect memory. Why would you ever use more tools to spend more tokens to construct a imperfect memory next to your session history?
I just don't get it.
ceejayoz · · focus · HN ↗
But that's one session. Isn't memory for… the next session?
unlikelytomato · · focus · HN ↗
dboreham · · focus · HN ↗
gregwebs · · focus · HN ↗
My thought though has always been that I don't want there to be agent-only designated documentation.
I use mattpocock/skills and that generates ADRs (Architectural Decision Records). That only uses skills, including a setup skill that will write a few pointers in AGENTS.md. I always have a CONTRIBUTING.md to document development flow and a CODING_STANDARDS.md. Between those and the README.md and architecture documentation and commit messages the agents seem to be able to find and use docs and keep them up to date. We are also writing a lot of specs and putting those in Github issues.
cyanydeez · · focus · HN ↗
Once qwen3.8-flash-next showed up, it cañ go "forever" with dynamic context pruning.
It's fascinating for local coding.
j45 · · focus · HN ↗
The slightest amount of guidance, input from experience can make a huge difference.
cyanydeez · · focus · HN ↗
And it's quite fascinating. Still dont think there's trillions of TAM out there if Qwen3.8-Flash-Next runs fast enough on $3k (before memory cartel) pricing.
girvo · · focus · HN ↗
And it’s an experimental likely undertrained model. Wait til Qwen 4 Flash…
cyanydeez · · focus · HN ↗
jon-wood · · focus · HN ↗
gregwebs · · focus · HN ↗
However, I found the same problems with agents that it has with humans (usually with humans it isn't TDD but code coverage requirements). The problem is the tests are now written to satisfy a bureaucracy rather than to properly verify the code. And there are real studies that point out the low value of these types of bureaucratic testing requirements. For example, when models are just required to write a test file, they don't produce a better result: <a href="https://arxiv.org/abs/2602.07900" rel="nofollow">https://arxiv.org/abs/2602.07900
I wrote a /verify skill that focuses on properly verifying code and I am finding it works a lot better. I do need to revise it now- in practice certain parts of the skill are doing all the heavy lifting and others are more dead weight. But it does seem to be properly orienting the agents towards finding defects. <a href="https://github.com/gregwebs/skills-sdlc/blob/main/skills/verify/SKILL.md" rel="nofollow">https://github.com/gregwebs/skills-sdlc/blob/main/skills/ver...
cyanydeez · · focus · HN ↗
Context then is: docs, code and tests. Rather than try to build some omniscient agent, the context is built where the agent needs it.
gregwebs · · focus · HN ↗
TDD doesn't help the agent figure out what it is supposed to do now- it is already writing the code and now and it must know what it is supposed to do to write it.
I am trying to use evidence-based approaches. I am only observing better results and tweaking now, but I plan to benchmark when tweaking is done.
cyanydeez · · focus · HN ↗
[dead]
cyanydeez · · focus · HN ↗
I'm working on javascript, and now I'm only doing this via typescript. I've successfully got this going:
1. Design a feature in plain language with the coding agen (opencode) and write a document for it.
2. Restart the context (or I use /compact to flush any errant details)
3. Pull the new plan into context and ask the agent to revise it follow Test Driven Development.
4. Depending on how big it is, the agent places it into multi stages, each with it's own document.
It then loops through the stages. It might be entirely based on the language you're using, but this loop seems strong enough.
Then when there's bugs, errors, anything, we update the doc, add more tests, then revise the code.
It's quite possible the harness you're using isn't setup properly. For me to do this locally, I had to hack on dynamic context pruning which I described here: <a href="https://news.ycombinator.com/item?id=49906637#49907641">https://news.ycombinator.com/item?id=49906637#49907641
gojogs · · focus · HN ↗
I write README and docs by hand, thus only important stuff goes in there. If it is not worth my attention to write down, it's not worth writing down.
Smaller context, easier to consume.
bushido · · focus · HN ↗
Essentially patterns the agents need to always think in. I also implemented a versioning system to the principles that need to be quoted in any comments which are there in the code. That way, when my principles evolve, so does the code.
I did package it up in a way that I can share it with friends [0]. System still evolving, but the last two-ish months that I've used it has served me really, really well, And it's been even better with the latest models.
I've had surprisingly good adherence from agents on this technique.
[0] <a href="https://principledriven.dev/" rel="nofollow">https://principledriven.dev/
skybrian · · focus · HN ↗
bushido · · focus · HN ↗
The github repo might be more helpful: <a href="https://github.com/Principle-Driven/pdd" rel="nofollow">https://github.com/Principle-Driven/pdd
skybrian · · focus · HN ↗
I was wondering about this bit:
> "Code comments cite a versioned token where code depends on the rule."
What does a token look like? I went looking for an example, but I don't know what I'm looking for.
bushido · · focus · HN ↗
visarga · · focus · HN ↗
The questions are designed specifically to identify the state the agent is in, in order to assign its steering. Changing this tool changes how the agent works. You just need to make the use of this tool a necessity so it does not forget to call it.
Unlike principles and memories, questions are more openended and tend sometimes to trigger the right mentality in the agent even before it gets to the policy itself. In my opinion agents get lost in local work losing the big picture, my questions jog the big picture back into attention.
1. call ssp tool, no arg, it just reads features from the repo, like most other tools -> ssp locates state and sends 4-5 questions
2. agent responds the first batch -> state is further narrowed down -> ssp sends a second batch of questions
3. agent responds again to the interview -> state gets finally pinned down -> agent gets the steering assigned
4. a log is created of this ssp session, and reflection on the log used to refine the questions and steerings
locknitpicker · · focus · HN ↗
It seems you tried to reinvent instruction files. What do you think is the difference between your approach and standard tools such as instruction files, AGENTS.md, and even skills?
mzhaase · · focus · HN ↗
bushido · · focus · HN ↗
If it helps, here's how you think about it, principles are not decisions, they're mental models and frameworks to allow making decisions. In my case, they're there to allow agents to make decisions without asking me. But the decisions cannot be recorded - the code serves that purpose. What can be recorded is why the commend for the block of code needs to cite what principle was used in the current shape, which is why there is a citation of the principle.
When the principle evolves, so does every decision, at the very least, every decision gets re-litigated to see if it needs to evolve.
PetriCasserole · · focus · HN ↗
Do you have any resources that guided your work?
bushido · · focus · HN ↗
A big part of it really comes to that I have led large product teams and platform engineering teams in the past, And I like giving feedback to my agents in the same way that I would plan things with my teams.
One of which is give them repeatable mental models which allow them to make decisions without relying on managers or me. The result is usually that the teams can work for weeks-to-months with autonomy.
I try achieving the same with my agents, so that they can work for 1-14 days without my intervention. I usually have them working on very very long running tasks/initiatives.
If you do want some resources, this is one of the pages that I've been using for nearly a decade to help people get started with mental models: <a href="https://fs.blog/mental-models/" rel="nofollow">https://fs.blog/mental-models/
vcryan · · focus · HN ↗
docheinestages · · focus · HN ↗
dboreham · · focus · HN ↗
isaachinman · · focus · HN ↗
<a href="https://github.com/isaachinman/encephalon" rel="nofollow">https://github.com/isaachinman/encephalon
mzhaase · · focus · HN ↗
bird0861 · · focus · HN ↗
But seriously, Opus has been garbage after 4.6 -- the public seems very easily fooled into thinking breadth == depth. Anthropic, I'll grant them, has been and is still an extraordinary data team. As for model alignment though... I have to wonder sometimes if they are even trying beyond just chat training.
"Guys, SI is just around the corner... Huh? What production DB? Anyway GPT2...I mean Mythos is too good it would be dangerous to release to the public!"
isaachinman · · focus · HN ↗
Neywiny · · focus · HN ↗
kolinko · · focus · HN ↗
You’re confusing, I think, memory system with llm finetuning. Completely different concepts.
Neywiny · · focus · HN ↗
YuechenLi · · focus · HN ↗
mcbuilder · · focus · HN ↗
dboreham · · focus · HN ↗
krishverma_2010 · · focus · HN ↗
[dead]
dboreham · · focus · HN ↗
kaydub · · focus · HN ↗
Every single project at work that has documentation/ADRs/etc ends up stale as fuck and it pollutes context more than it helps. Even the agent generated docs. Hell, especially the agent generated docs. It's like the LLMs can't just DELETE something, they always amend. So now there's stale details in there about something we abandoned polluting context.
The code is the documentation, none of this stuff is needed. I'm seeing WAY too many devs/engineers go down the rabbit hole of trying to build these tools and systems but not actually getting meaningful work done.
I've been doing my best to remove and delete all the documentation. I usually keep a super high level AGENTS.md/CLAUDE.md file and that's it. The llm/agent can figure it out. You might argue I waste 5 minutes and a dollar or two for the agent to "figure it out" each time, but I'm pretty certain people are wasting FAR more time and money building these things out and maintaining them in each of their projects (or NOT maintaining them and they're actively BAD for their projects).
relaxfolio · · focus · HN ↗
[dead]
spike021 · · focus · HN ↗
If I say "use jq instead of writing a python script to parse json" it should never write adhoc python scripts to parse json. Yet that constantly happens to me anyway.
ACCount39 · · focus · HN ↗
ozim · · focus · HN ↗
gojogs · · focus · HN ↗
zahrevsky · · focus · HN ↗
aliasxneo · · focus · HN ↗
gf000 · · focus · HN ↗
jon-wood · · focus · HN ↗
PcChip · · focus · HN ↗
rmunn · · focus · HN ↗
spike021 · · focus · HN ↗
It's not like it's being rewritten for efficiency. "Just because".
amelius · · focus · HN ↗
girvo · · focus · HN ↗
dboreham · · focus · HN ↗
locknitpicker · · focus · HN ↗
You haven't been paying attention then. I routinely see Claude and GPT models churning out python code to do stupid things like linting. Last week I even had a TypeScript project with prettier configured all over the place, including in a custom agent skill I added with the express purpose of getting the damned model to lint the code, being constantly prompted to run ad-hoc python code supposedly to format whitespaces. I even explicitly prompted one session to just use npm run lint, where I pointed out the exact line of code where prettier was invoked, and the session still churned python code to hande whitespaces.
yulaow · · focus · HN ↗
jkhdigital · · focus · HN ↗
koolba · · focus · HN ↗
Otherwise you can get a python one liner that execs a different script engine.
devmor · · focus · HN ↗
TZubiri · · focus · HN ↗
Terr_ · · focus · HN ↗
altmanaltman · · focus · HN ↗
user_of_the_wek · · focus · HN ↗
sick_of_slop · · focus · HN ↗
[dead]
locknitpicker · · focus · HN ↗
I think there is a deeper problem emerging from this sort of behavior. Even when we bother to create agent skills with there own scripts that call tools like jq a specific way to achieve a goal, AI coding assistants and agents still go way out of their way to generate ad-hoc scripts to do the most absurdly stupid tasks such as parsing output in structured language, and even remove whitespaces from a markdown file. This means AI coding assistants and coding agents treat agent skills as mere suggestions of using a alternative option that more often than not the choose to ignore.
This has a very dangerous implication: your average user is trained to develop a pavlovian reflex to authorize agents to just execute their ad-hoc scripting code with our own permissions and credentials in our systems, which includes the ability to call anything over the internet.
rojaneerdev · · focus · HN ↗
[dead]
ozim · · focus · HN ↗
<a href="https://gist.github.com/cynthiateeters/6868ca26c059a3106cd93736d33015af" rel="nofollow">https://gist.github.com/cynthiateeters/6868ca26c059a3106cd93...
HisashiSpace · · focus · HN ↗
[dead]
fennect · · focus · HN ↗
I’m the author of the project.
triyambakam · · focus · HN ↗
joquarky · · focus · HN ↗
eivindmeyer · · focus · HN ↗
[dead]
kadhirvelm · · focus · HN ↗
greatergoodguy · · focus · HN ↗
Since then I've gone down the rabbit hole of really digging deep into current AI research and especially what AI whistleblowers are currently saying. And I can't even express how existentially scared shitless I am.
mewpmewp2 · · focus · HN ↗
dingaling911 · · focus · HN ↗
altcognito · · focus · HN ↗
That being said, the danger is measured in computer damage, which can be a lot personally and to a company, but less existential, so your mileage may vary as to how "scary" it is.
Xirdus · · focus · HN ↗
throwaway27448 · · focus · HN ↗
Forgeties79 · · focus · HN ↗
apsurd · · focus · HN ↗
Forgeties79 · · focus · HN ↗
iamwil · · focus · HN ↗
zahrevsky · · focus · HN ↗
> No one rewatches a team meeting from 3 years ago to remember constraints around a feature. People write things down and use those records instead.
Memory plugins don't re-read old transcripts. Memory snippets are basically the notes that people write down after the meeting.
The problem, however, is different. The problem with memory plugins is that your agent basically writes a note every time someone says a sentence, and then tries to work with those 5000 notes.
Instead, the agent should recognise what's important and write only that. And that is, of course, a documentation. (And ADRs, if you want not only a description of the final state, but also the trajectory of how we arrived to it. Which, arguably, contains more information than the docs alone.)
Of course, another difference is that notes are immutable, append-only and don't have a lot of structure. This is of course done to be able to store lots and lots of notes: they should be independent. The main problem is that increasing the number of notes adds not enough benefits to compensate for downsides of this structure-less immutable format.
popalchemist · · focus · HN ↗
nijave · · focus · HN ↗
pornel · · focus · HN ↗
I've been bitten by agent-written ADRs. Agents carelessly add extrapolated details and speculated nice-to-haves that I never asked for, and this becomes a source of bloat that keeps coming back like a boomerang.
zahrevsky · · focus · HN ↗
sathish316 · · focus · HN ↗
Lack of this capability makes automated principles or patterns update a recipe for more bloat.
otterley · · focus · HN ↗
jen729w · · focus · HN ↗
`cd` to a folder. Launch `claude`. Do your work. Save scripts and documentation in that folder. `/resume` previous conversations from that folder.
That's it. That's the trick.
Now, having very static, very well-defined folders helps a lot. I'm Johnny.Decimal so I have numbered folders for everything I do. So my process when I want to use my 'process a travel booking from my email to my calendar' script is:
- `jd tripsy`
- `claude`- Say 'hey Claude, there's a new email in my inbox please'.
- Done.
nijave · · focus · HN ↗
Legally and culinarily, they're vegetables. Botanically they're fruit.
Tomatoes are members of the nightshare family which includes tobacco, potato, and chili peppers.
They are used in popular recipes like Mexican salsa and Italian pasta sauce. Pasta sauce commonly uses the San Marzano variety. Here's a <picture> of San Marzano tomatoes from our Italy trip.
Rhett likes tomatoes in all dishes. Link only likes tomatoes if they're blended in a dish like pasta sauce or tomato soup. Frank is allergic and can't have tomatoes.
Tomato prices are up due to a bad <year> yield in <country>.
---
This is all memory. Where does it go?
jen729w · · focus · HN ↗
What is it that you are doing with tomatoes today? Store that locally. Sounds like you're cooking. I put those here, then I can find them again really easily.
mcapodici · · focus · HN ↗
If I want the LLM to remember something I ask it to update some docs, and even check that docs are consistent across the board after doing so.
The only thing I want the LLM to remember everywhere is talk like a human (no load-bearing, not this/that etc...), so I have an AGENT.md for that.
jsemrau · · focus · HN ↗
folayii · · focus · HN ↗
[dead]
[deleted] · · focus · HN ↗
[deleted]
chaostheory · · focus · HN ↗
jmtulloss · · focus · HN ↗
Snarky comment aside, I am very interested in how we evaluate the performance of these systems and what kinds of work match best with different approaches.
lewelove · · focus · HN ↗
<a href="https://xkcd.com/927/" rel="nofollow">https://xkcd.com/927/
favori995749721 · · focus · HN ↗
[dead]
ramoz · · focus · HN ↗
<a href="https://backnotprop.com/blog/context-monorepos/" rel="nofollow">https://backnotprop.com/blog/context-monorepos/
When I ask "what happened to x?" ... the agent has everything it needs to give me that answer. When it plans the next feature it can validate assumptions against previous decions made in my `decisions` folder.kaydub · · focus · HN ↗
I don't think you need more documentation or these memory systems. The code IS the documentation. Lots of this stuff is unnecessary. It's a bunch of people doing their special rain dances and then when it happens to rain they say "I did that"
nicwolff · · focus · HN ↗
<a href="https://cline.bot/blog/memory-bank-how-to-make-cline-an-ai-agent-that-never-forgets" rel="nofollow">https://cline.bot/blog/memory-bank-how-to-make-cline-an-ai-a...
minimaxa · · focus · HN ↗
[dead]
stbenjam · · focus · HN ↗
[dead]
DriverDaily · · focus · HN ↗
Like, you can quickly lookup related ideas based on what came before and after, causes, effects, just like calling relationships a graph database.
Documents can’t be queried efficiently like that, you need a database.
apsurd · · focus · HN ↗
I take the article's point more directly. It's just a straighter line to have clear communication through documentation than to fuddle around with the perfect memory setup.
Garlef · · focus · HN ↗
I tried an approach based on the following idea recently and it's amazing - Lint rules where the error messages contain an explanation on how to deal with the issue.
<a href="https://habit-hooks.com/" rel="nofollow">https://habit-hooks.com/
I'm using it to foster IOSP (integration operation segregation principle) for example.
jghn · · focus · HN ↗
Garlef · · focus · HN ↗
<a href="https://youtu.be/6AgndHSkHFI?t=238" rel="nofollow">https://youtu.be/6AgndHSkHFI?t=238
jcjmcclean · · focus · HN ↗
Then for example you could write your own hook and convert existing code smell documentation which agents ignore into a format that works with the hook.
Garlef · · focus · HN ↗
That's what I did; I did not use the library I linked to ~ It served only as an inspiration.
Instead, I let the agents create custom lint rules (using eslint, pylint, ...) and add custom coaching error messages based on where I want to take my codebase.
jcjmcclean · · focus · HN ↗
spacebanana7 · · focus · HN ↗
fxtentacle · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
DelightOne · · focus · HN ↗
Or is that the secret sauce no one wants to share
fxtentacle · · focus · HN ↗
But some of the resellers have freebies, like:
<a href="https://app.primeintellect.ai/dashboard/environments?ex_sort=by_sections" rel="nofollow">https://app.primeintellect.ai/dashboard/environments?ex_sort...
<a href="https://github.com/sierra-research/tau2-bench" rel="nofollow">https://github.com/sierra-research/tau2-bench
balder1991 · · focus · HN ↗
OJFord · · focus · HN ↗
szundi · · focus · HN ↗
[dead]
Garlef · · focus · HN ↗
Just asking for small functions will make the agent write small functions ~ but not in a good way (for example they just cut `doOneThingAndTheOther` in half and call the second half `doOneThingAndTheOther2`)
tankiya · · focus · HN ↗
[dead]
suki_buildz · · focus · HN ↗
[dead]
jason1cho · · focus · HN ↗
Indeed, AI fanatics lack critical thinking. They easily accept the idea that AI needs documentation, but merely disagree the method to approach it.
indrex · · focus · HN ↗
fullstackwife · · focus · HN ↗
[dead]
ContinuityLab · · focus · HN ↗
locknitpicker · · focus · HN ↗
Isn't memory just documentation that agents create and update on the fly to fill in the documentation hole that your average user leaves open?
There was a time when the need to provide context and signal was a key topic in LLM and AI-assisted coding. People talked about MCPs AGENTS.md and README.md and agent skills and even comments, descriptive tests, and naming conventions. Supposedly the theory was that if you provide context, agents wouldn't misbehave so much. Then plan mode and spec-kit approaches stepped in with approaches aimed at providing that context when starting sessions, because users never bothered with docs or MCPs or AGENTS.md or anything. But then plan mode and spec-kit approaches were considered too laborious and requiring too much from users. Then, because users systematically failed to provide context, this task was finally given to agents. And things just worked.
You should ask yourself why in software engineering circles documentation is seen more as a problem than a solution to any problem.
yethikrishna · · focus · HN ↗
[dead]
atworkc · · focus · HN ↗
The only "prompt" is in AGENTS.md saying that this thing exists and there's a map.md <- which is a one liner reference to whatever the agent stores in there.
And usually, I tackle a new feature, and at some point tell it to store to jot down notes in workbench if I'm comfortable with it (and if it needs to be stored in memory)
This makes it easy to just spawn other agents and such from a good point, I just point them that stuff is in workbench.
Token costs, seem decent and I can always just delete stuff in there as its purpose is ephemeral.
ArtRichards · · focus · HN ↗
<a href="https://github.com/ArtRichards/docs-cli" rel="nofollow">https://github.com/ArtRichards/docs-cli
and
<a href="https://artrichards.github.io/agent-playbook-suite/blog/" rel="nofollow">https://artrichards.github.io/agent-playbook-suite/blog/
chrisweekly · · focus · HN ↗
hedgehog · · focus · HN ↗
aniceperson · · focus · HN ↗
engeenie · · focus · HN ↗
[dead]
sreekanth850 · · focus · HN ↗
[dead]
b-karl · · focus · HN ↗
We use a private Claude plugin marketplace for internal plugins and skills and I try to regularly prune my memory, migrating relevant stuff to a proper home (skills in the marketplace, docs in repos or Notion etc) and prune outdated information.
jon-wood · · focus · HN ↗
OhNoNotAgain_99 · · focus · HN ↗
[dead]
bob1029 · · focus · HN ↗
The idea of dumping everything into a big database and hoping the model will write the correct queries does work out to some extent. It's a very enchanting idea. However, it pales in comparison to having a dedicated tool per type. The outcomes seem to be much better when joins between low cardinality types occur within the token stream.
If your agent does need access to some enterprise knowledge base, I would give it a lexical search capability and not overthink it with vector shenanigans.
Tools are the only thing you need if you build them right. I don't even have a system prompt anymore aside from injecting the name of the robot and the current user's name. Keep in mind that all aspects of tools can be dynamic over time. I've got some where the description is composed by hundreds of lines of conditional string builder depending on the current state of the conversation.
skinfaxi · · focus · HN ↗
simianwords · · focus · HN ↗
I think what you are trying to say is that “give the LLMs access to information and it can figure out how to use it. Don’t get fancy about how to structure the data”.
If that’s the case it has largely turned out to be true. RAGs have fallen out of fashion.
Where does that put AGENTS.md though? Is it worth spending time to structure it or just dump it.
voiper1 · · focus · HN ↗
I've worked with docs that get _updated_, and progressive disclosure: make references to more obscure features in their own page, so they don't majorly bloat the context.
pawel_nowak · · focus · HN ↗
[dead]
derin-picment · · focus · HN ↗
[dead]
scotttaylor · · focus · HN ↗
[dead]
scotttaylor · · focus · HN ↗
[dead]
jew-and-proud · · focus · HN ↗
[dead]
firemelt · · focus · HN ↗
zenapollo · · focus · HN ↗
.agents/plans/<plan-name>/
.agents/notes/<topic>/
.agents/knowledge/<topic>/
I only commit knowledge and if knowledge gets big i add .agents/knowledge/INDEX.md
This is a good mix of human readable and agent fluent. Notes are ephemeral, knowledge is permanent.
I have a rule for knowledge that it has to be stable and mostly permanent (though updatable). And the agents are not allowed to post there unless docs are clean organized and with permission.
Notes are for jotting things down and handoffs, massaging a featureset. Agent can document at will.
Still WIP.
sdevonoes · · focus · HN ↗
kursus · · focus · HN ↗
stuartd · · focus · HN ↗
tw1984 · · focus · HN ↗
balder1991 · · focus · HN ↗
clickety_clack · · focus · HN ↗
byteknight · · focus · HN ↗
notjes · · focus · HN ↗
lmeyerov · · focus · HN ↗
[dead]
koct9i · · focus · HN ↗
k__ · · focus · HN ↗
nahid-fahh · · focus · HN ↗
[dead]
marvstazar · · focus · HN ↗
Organizing the markdown files by feature/issue/change makes it easier for the LLM to search for the appropriate documentation. Coupled with well-broken down Claude rules files, and Claude Code (or other harnesses) get better as you make more changes.
dhruv_dube · · focus · HN ↗
[dead]
katanascreener · · focus · HN ↗
[dead]
WillAdams · · focus · HN ↗
<a href="https://news.ycombinator.com/item?id=47300747">https://news.ycombinator.com/item?id=47300747
"We should revisit literate programming in the agent era" (silly.business) 292 points, 251 comments
muaguy · · focus · HN ↗
[dead]
openamer · · focus · HN ↗
[dead]
dham · · focus · HN ↗
jrochkind1 · · focus · HN ↗
CapitalistCartr · · focus · HN ↗
This is a first draft; his github is better than his article. Looking through it, Consult actually works. The agent doesn't pick documents blind. Every scope has a catalog file that describes each document: what it covers, when to open it. These catalogs seem to load in to the start of each session, so the agent gets a little map without reading every file. Code navigation seems the same. Each index document has a short description and a "read_if", and subindexes are opened when their condition matches the job. This looks pretty well laid out, which I would never have guessed from the article.
aaronscott · · focus · HN ↗
This way the primary agent only has relevant information in their context to make decisions and take actions.
Context management is still under valued imo.
CapitalistCartr · · focus · HN ↗
themgt · · focus · HN ↗
mikeryan · · focus · HN ↗
Not today. IMHO Documentation is just a form of structured memory and it’s all just context. Getting that context right is a hard problem and there’s a lot of different ways to skin that cat.
chapmancl · · focus · HN ↗
[dead]
crazygringo · · focus · HN ↗
I don't know, but I have a pretty standard (I think?) setup, and Claude manages to find every relevant file every time. But I've also only used Claude for greenfield projects, where "documentation is primary, and code flows from documentation" is the philosophy.
I have CLAUDE.md describe all the types of documentation files and the directory structure. And then Claude is pretty aggressive (automatically) about always inserting cross-references everywhere. So a feature description will reference the ADR's that it implements, the ADR's say what feature implements them. A code file will make reference to the "implementation design" document that describes the motivation behind which iOS elements were chosen, how the animation is defined in a particular way that doesn't break another animation, and so forth. So I've really never run into a situation where Claude failed to read a file it should have. I've been pleasantly surprised.
I would say that the one really big thing I've had to learn is to teach Claude both in CLAUDE.md and in the header of every top-level design document, that keeping documentation current and in sync is paramount. Because its default seems to be to keep history and append, e.g. by default it will take a section of a document and mark it "[DEPRECATED]" and add the new version below. So my instructions are pretty clear in having it always be aggressive in maintaining current state only, always replace rather than append. And if there's anything we want to save from the previous approach (e.g. we did it X way previously and it failed because Y), then just add that as a new short note in the new current-state text, possibly with a pointer to a commit or tag or something.
So this seems to solve both recall and staleness in my projects at least.
The only thing I still haven't found a solution for is numbering. Claude is always giving everything numbers, like F23 for feature 23. But I'm always changing the order of things, inserting new things, deleting things, so I wind up with a sequence of development work that goes in order like "Phase 9", "Phase 9b", "Phase 9e", "Phase 11", "Phase 12". I'm halfway ready to abandon numbers entirely and just start giving things names from noun collections instead, so every feature is named after an animal, every ADR is named after a kitchen implement, or something. Or just four-digit hex codes chosen at random. Curious if anyone else has found what works.
hedgehog · · focus · HN ↗
demibabs · · focus · HN ↗
athrowaway3z · · focus · HN ↗
When you have knowledge distributed in markdown files; finding them puts the path/filename into context as well as some indication of document size. (If its on line 1200 or line 20). This is extremely valuable for picking what ought to be focused on next.
RAG on the other hand creates the hardest challenge for these models. It instead puts 5 ideas with the highest similarity into the context in full.
Its the difference between having to remember a set of numbers when in a crowd that's talking about stuff, and having to remember them when the crowd is shouting out random numbers. The similarity in the task makes things harder. SoTA models work despite this, but its extra-gambling while you're already gambling.
panarky · · focus · HN ↗
In this implementation, Markdown should be considered harmful.
Operator Memory injects `.operator-shared/operator.md` and `.operator-shared/index/.md` directly into your agent's instructions before you even write the first prompt.
So if you clone a repo or review a PR where a bad actor put malicious instructions in these files, now your agent executes those instructions automatically and silently.
It could exfil `.env` and `~/.ssh/`, change `~/.bashrc`, all kinds of dirty deeds.
Agents are pretty good now about not running prompt injections hidden in code and Markdown, but this plugin bypasses all of that, and puts the prompt injection right in the system prompt.
And with higher priority than AGENTS.md and CLAUDE.md.
Seems bad.
yesb · · focus · HN ↗
I think what you're observing is that there is more to information retrieval i.e. "retrieval" in RAG than slapping everything into a vector database and calling it a day. There's no such requirement in RAG to mindlessly load the k nearest neighbors into your context and see what happens. That's a very rudimentary implementation.
This markdown system I'd argue is RAG as well. You're just doing the retrieval in a way customized for the problem at hand. If you have a precise method of retrieving the most relevant things, obviously use that rather than a similarity metric. If I'm reading correctly, this markdown system is basically a knowledge graph which is not a new idea.
zhuangweiguo · · focus · HN ↗
[dead]
nahid-fahh · · focus · HN ↗
[dead]
haroldopina · · focus · HN ↗
[dead]
tabbybyte · · focus · HN ↗
tabbybyte · · focus · HN ↗
zx8080 · · focus · HN ↗
cracell · · focus · HN ↗
scotty79 · · focus · HN ↗
baalimago · · focus · HN ↗
I built <a href="https://github.com/baalimago/slivingdoc" rel="nofollow">https://github.com/baalimago/slivingdoc which achieves the same thing as OP's tool, but with s3 as backend (or hosted, for the lazy one with $1/month to spare). But it's simpler, the agents simply commit or pull using git semantics to allow them to organize their memories as they best see fit. Using text is the way to go, but how the information is encoded is likely never going to be "solved", and will likely differ agent per agent.
ravenstine · · focus · HN ↗
crazygringo · · focus · HN ↗
"Code is the documentation" doesn't solve the problem in my experience. Because what happens is that you still need a lot of "why" comments in the code, and then these go stale, so you still have the same problem you have to solve. And so I find that markdown documentation is a lot easier to organize and review in a structured hierarchical way in one place, than code comments sprinkled across the repo.
nijave · · focus · HN ↗
It also doesn't seem to have a way to separate preference from factual memory which is useful when you have various humans interacting with the same agent. One human might prefer a certain output style over the other and that's something a memory framework can also address.
Like the MCP articles a few months ago, it also seems to assume all agents are cli coding harnesses running on your local machine. We have a handful of other things like chat bots, event-driven agents running on servers, chat driven agents running on servers in sandboxes--there's not a single filesystem and even if there were one, having multiple agents try to edit it at once would corrupt it.
chapmancl · · focus · HN ↗
[dead]
higeorge13 · · focus · HN ↗
[dead]
AllegedAlec · · focus · HN ↗
It doesn't even understand when its current context isn't...
See the whole issue of refining a concept, removing a few points or notions, and it will write a paragraph why a point that isn't in the text anymore when no version a human will read ever had that point in it.
Clumps4Linux · · focus · HN ↗
[dead]
charles_f · · focus · HN ↗
jookers · · focus · HN ↗
<a href="https://gist.github.com/karpathy/442a6bf555914893e9891c11519de94f" rel="nofollow">https://gist.github.com/karpathy/442a6bf555914893e9891c11519...
rogeliodh · · focus · HN ↗
Almey · · focus · HN ↗
Almey · · focus · HN ↗
<a href="https://github.com/Clumps4Linux/SEP/blob/main/SEP-1.1.1.txt" rel="nofollow">https://github.com/Clumps4Linux/SEP/blob/main/SEP-1.1.1.txt
kaydub · · focus · HN ↗
All this stuff is LLM rube goldberg machines. It just pollutes context.
I barely use AGENTS.md/CLAUDE.md these days. And where they remain, it's super basic high level stuff.
I'm honestly still kicking myself in the ass on many projects where I did something similar to this. I kept tons of markdown docs and decision docs. Now those things are just causing problems because they got stale. Even after having sessions of reconciling documentation, the LLM just gets confused.
le-mark · · focus · HN ↗
kaydub · · focus · HN ↗
AppleBananaPie · · focus · HN ↗
I think it's going to be an incredibly common, maybe universal cycle people will go through working with AI until they realize it doesn't work long term.
smokel · · focus · HN ↗
Why something exists, and how it connects to the outside world may be documented in comments, but more often than not it isn't.
haukebri · · focus · HN ↗
[dead]
arcanemachiner · · focus · HN ↗
kaydub · · focus · HN ↗
smokel · · focus · HN ↗
Consider working on software for a coffee machine. Why is pin 42 (GRIND) activated every now and then? Would you really want to document that in Git commit messages?
kaydub · · focus · HN ↗
I'm seeing decision files that are so big the LLM can't fit it all in context on some projects. Then the LLM makes decisions that revert previous ones and later sessions don't pick that up so it sticks to the original decision. Now in some sessions, every so often I have to remind the LLM, "no, we changed that later, we do it this way now"
And I'm seeing our knowledge base grow to a completely useless giant mess of stale, outdated, duplicated, or superfluous info. LLMs often pull unrelated info or confuse similar but different things or get old documentation for something that's been updated to new documentation in a different part of the knowledge base. And these are LLMs generating the docs. And we have LLMs and agents reconciling. But it doesn't seem to always get everything.
For code comments, it's terrible because the comments are starting to get larger than the code. A large chunk of the comment can be discerned from the code itself. Then the comment has details on why that maybe don't quite make much sense. It's like the LLMs start using words in a specific context that doesn't really apply to the word in normal spoken english. Then it will also often include a specific JIRA ticket id, you check the JIRA ticket, you see that yeah, the code was changed because of that JIRA ticket, but it's not really related to the ticket itself, it was just a blocker. But now the comment forever links it to THAT ticket (And then now sometimes the LLM pulls in that ticket with the atlassian mcp).
haukebri · · focus · HN ↗
[dead]
rectang · · focus · HN ↗
But this article argues that LLMs do better when the context is smaller — when it can understand the totality of the task with as little context as possible. And so having correct API-level docs is greatly advantageous. Anecdotally, this rings true to me — when the local context is good and clear, the LLM writes code matching my intent even when my prompt is sloppy and poorly specified.
Rejoice! The LLM will write the docs for you, relieving you of most of the work.
However without intervention, it will do too much and record absurdly verbose docs (similar to how an LLM will relentlessly refactor your code until you instruct it to move in minimal, incremental changesets). You will still need to edit down what the LLM generates.
klodolph · · focus · HN ↗
trinsic2 · · focus · HN ↗
kaydub · · focus · HN ↗
The docs end up stale or you waste a ton of time and tokens keeping them up to date. And you can't even trust the LLMs to keep the docs up to date, you will still need to review it yourself keep a bunch of shit out of them.
joquarky · · focus · HN ↗
kaydub · · focus · HN ↗
I don't mind SOME documentation. I'm growing extremely frustrated with the absolute MOUNTAIN of docs, comments, decision documents, etc being created. I'm so fucking tired of the LLM rube goldberg machines being built.
kaydub · · focus · HN ↗
The docs just make you feel good. They're not worth anything. They just create more work if anything because now you don't only have to fight entropy in your codebase but also your docs.
rectang · · focus · HN ↗
kaydub · · focus · HN ↗
I'm not the one bolting things on claiming it increases productivity. Why would I ADD stuff without any proof or evidence that it works? I'm simply NOT adding these things.
JohnBooty · · focus · HN ↗
Notice the things they struggle with and the small problems (typically, environment issues IME) they repeatedly encounter and re-solve across multiple sessions. That is what your instructions should cover. When possible, move those instructions into skills, so they get loaded into context selectively instead of on every session. (Example: instructions for running specs, placed into a skill that only gets loaded into context when it’s time to run specs)
Again, this can be automated by the LLMs themselves: both Codex and Claude (and I’m assuming other major harnesses) know how to read their own transcripts and are good
kaydub · · focus · HN ↗
The LLMs will always have these suggestions on things that will "reduce friction in the future" but I really think the LLM is a sycophant glazing you. Because even when you add those skills or prompts at some point the LLM ignores it or does whatever it wants to do anyways. Better to leave most of that out and deal with it when it arises.
Even your example, instructions for running specs... none of the current frontier models need this at all.
joquarky · · focus · HN ↗
kaydub · · focus · HN ↗
At work it's all Anthropic. I spent a while on Sonnet models but have been pretty consistent about using Opus now. For personal projects I bounce between Sol/Terra and Gemini Flash.
Most of the models are good enough now. You don't need all these documents, memory, plugins, skills, etc. MCPs are good for hooking up to external sources of data (at work we have gitlab mcp, datadog mcp, and our knowledge base mcp... which the knowledge base is mostly worthless now because of all this AI generated documentation, but I digress). I DO still use beads on a lot of projects, but not always. My prompts these days are lazy as fuck: "review this repo, check out this other repo with architectural patterns we should follow and libraries/terraform/etc we should use, <basic desc of goal>. What do you think? Let's review everything and discuss before we start building"
Oh no, I might have to stop the llm agent on occasion, or review what they did and tell them to fix a couple things. Way better use of tokens and my time than creating some rube goldberg machine. Way better than keeping a giant decision file that goes stale and causes context corruption because the llm only saw the OLD decision and not the UPDATED decision.
JohnBooty · · focus · HN ↗
That must be a remarkable model. Smart enough to flawlessly discover everything on its own in every session... but also it just straight up just doesn't obey direction.
Hmmm.
kaydub · · focus · HN ↗
I didn't say give no direction.
I'm saying a lot of devs/engineers little rain dances aren't really bringing the rain.
Keep a basic CLAUDE.md/AGENTS.md that keeps high level details. Have a way to reference other projects/apps/codebases to use as reference architecture.
Don't keep tons of markdown files. Don't keep "decision" documents (these are probably the worst). Don't make skills for every little thing (superpowers are dead these days, stop using them). Oh yeah, don't have the LLM generate docs for reference later... if it can generate the docs it doesn't really need them.
JohnBooty · · focus · HN ↗
Uncharitably, I think your claim is absolute junk. In nontrivial projects there are codebase conventions and environment specifics that are not self-obvious and cannot be discovered without significant effort or perhaps at all.
Even if we generously assume that the LLM is smart enough to discover these things flawlessly, that still is repeated work that needs to be repeated in every session, and that repeated work takes at least an order of magnitude more tokens than simply giving them concise directions.
Apologies for being unclear. They don't literally need directions to run specs. I was giving an example of how context-specific directions should be given via skills or similar, so they don't pollute context on every session.If you want a real-world example, our app spins up a different docker image per git worktree. Agent-browser needs to know the app's URL, obivously, to perform automated QA. The URL will differ per worktree. It gets the URL from a script in bin/. It could rediscover this on its own in every single session, or we could tell it how.
We've also found they follow codebase conventions much more reliably if informed of them rather than inferring them. For example, we prefer lean models. Cross-model concerns go in service objects. Even if LLMs could discover this flawlessly, it would require a lot of work in every single session to re-learn it -- re-grepping the app, doing a lot of LSP querying, etc. There's just no version of reality where that doesn't burn a lot of tokens doing the same work over and over again in every session.
It's frightening that you're apparently using some kind of unreleased magical model that not only needs zero direction, but also will not even obey directions.kaydub · · focus · HN ↗
You're just stuck on your dogma.
JohnBooty · · focus · HN ↗
However. At least in the projects I've worked on, it's trivial to read the transcripts and notice that without enough direction, there are certain things the agents will burn lots of tokens to rediscover in every session. That's absolutely not dogma.
kaydub · · focus · HN ↗
And I think a lot of places are burning way more on the latter than the former. I think the former is more cost effective due to caching as well. So I HEAVILY lean towards the former especially with how effective I think it is these days.
kaydub · · focus · HN ↗
icedchai · · focus · HN ↗
kaydub · · focus · HN ↗
I don't think I've ever seen an agent review a doc and then NOT also go look at the code. So skip the middleman, just have the agent look at the code.
joquarky · · focus · HN ↗
kaydub · · focus · HN ↗
NBJack · · focus · HN ↗
ramesh31 · · focus · HN ↗
enraged_camel · · focus · HN ↗
We have heard this nonsense from the "we don't need to write comments, code should be self-documenting" types for decades. It was wrong in that context, and it is wrong in this one.
Code tells you how a system works. It does not tell you why it works that way. That is what memory is for. It exists so that your AI does not keep undoing past decisions when it writes or refactors code.
kaydub · · focus · HN ↗
Old decisions always end up in the docs. If the LLM gets a whiff of an old decision, but doesn't get the update to that decision, well now you're doing things the old way again.
dregitsky · · focus · HN ↗
But not sure it works in an age where most code is LLM-generated. Especially if that code is not even reviewed by humans (irresponsible or not, it's happening), and commit messages are also generated by AI. I think something is needed to separate "what did the human operator intend" from what the agent went and built.
I do agree that this gets way overengineered. My approach has been more or less what you stopped doing though - committing all our timestamped "plan/implementation docs" and "investigation docs" that document what the user wanted + empirical findings, and making all prior session transcripts searchable. It's seemed mostly helpful? For whatever reason I haven't run into many staleness problems so far.
kaydub · · focus · HN ↗
I'm mainly aiming my frustrations at all the markdown files being committed, all the additions to knowledge bases, all the comments in the code (especially the ones referencing specific JIRA tickets). This stuff isn't helpful, it gets hella stale. I've had the LLM fuck up plenty due to these docs and comments.
Commit message ARE EXACTLY where architectural decisions or nuance should go. Not another fucking .md or more comments.
jjfoooo4 · · focus · HN ↗
kaydub · · focus · HN ↗
Make a high level .md, let it rip, iterate. You can give specifics in your original prompt.
I don't need to document the framework or the libraries etc. Pre-code, I tell it in the prompt one time. After it inits the project it's in the code.
mmcnl · · focus · HN ↗
kaydub · · focus · HN ↗
If it's not, that's what commit messages are for.
hosh · · focus · HN ↗
When it is not, it has to be reasoned out and does not work well for documentation.
Other things that code and testes do not capture well:
- promises (as in Promise Theory) made to other parties
- Constraints-inducing-properites, as in Roy Fielding / Christopher Alexander. While tests, and property testing can capture properties, there is no formal connection to the constraints that induces fhem. By constraints, I am not talking about business requirements and business value — those are better understood through Promise Theory. I am talking about things like at-least-once delivery or total ordering (from append-only constraint)
- grammars, as in pattern panguages (not just patterns) a la Alexander / Fielding are also not captured in code alone. These tell both humans ans AI how to extend a pattern, and how to identify anti-patterns (when they violate a constraint-inducing-property)
- LLMs are trained with many different worldviews and bounded contexts at the same time, and is very capable of translating across it. However, these need to be soelled out, otherwise it would talk in whatever it infers
Specifications written for the exact way components are wited together run into that stale doc problem. Although it takes much more human attention and token burn to describe things in terms of pattern language and promise theory, it becomes easier over time. The actual implementation plan tends to fall out more cleaning when all those other stuff are at least considered. This is where I have been spending most of my time.
chaostheory · · focus · HN ↗
kaydub · · focus · HN ↗
01100011 · · focus · HN ↗
In my codebase it is difficult to get agreement on comments and documentation so rather than rely on it I adapted. One of the first things I did when I succumbed to agentic development was to point codex at the code and ask it to generate a high level description of where important files, such as our public API, reside, what the hierarchy is, what the code does, etc. In my case, this level of documentation is fairly static if I avoid implementation details. So now I have a handful of agent files in my tree and it seems to save quite a few tokens and improve my results. I frequently have other devs ask me how I get such good results when doing agentic reviews of their changes(always my first step now before I start my human review). I also include instructions in the agents files instructing the agent to maintain the agent files if any relevant changes are made. It seems to work quite well for me.
kaydub · · focus · HN ↗
If you got the LLM to generate the docs, they don't need the docs.
mgfist · · focus · HN ↗
All code is written under constraints, and most constraints live outside the code.
kaydub · · focus · HN ↗
When it's not, there are commit messages.
Please for the love of god, quit generating markdown files (especially having the LLM generate the file, because if it could generate it, it doesn't need it), quit generating "decision" docs, and stop having more comments than code.
mike-akdeniz · · focus · HN ↗
[dead]
ChimpWithHat · · focus · HN ↗
kaydub · · focus · HN ↗
It's a LOT of cargo-culting overly complex AI workflows. It's devs/engineers making rube goldberg machines.
Everyone is doing their own little rain dance and when it rains they say "see, I told you it works"
shinokami · · focus · HN ↗
for a recent long-running task in a large repo, I used a tiny markdown spec with just the functional requirements and some context
no OpenSpec/Spec Kit or anything around it, just one file
and it worked really well as a temporary source of truth across different agent sessions
JohnBooty · · focus · HN ↗
Trivial example: LLMs struggled with QA on our app. They didn’t know how to find the seeded test accounts. They would create new ones and do it wrong, or find the seeded test accounts but not know their passwords because they were encrypted, so they’d change the passwords but not tell the other agents. Shitloads of tokens burned. It was a no-brainer to just give them the credentials in a skill that gets loaded when they do QA.
You could say that’s an environment issue, not a code issue. But as far as actual code IME at a minimum we have to tell the agents our general repo structure and architecture patterns.
The days of “you are the world’s greatest Python programmer, write good code” or whatever are over (if that style of prompting even worked in the first place) and as I said less is more. But, still….
kaydub · · focus · HN ↗
Yeah, that's exactly what I'd say. Or maybe you're just approaching the problem now.
You have to prompt these QA agents, correct? Why not give the instructions on where to get credentials in the original prompt?
> But as far as actual code IME at a minimum we have to tell the agents our general repo structure and architecture patterns.
I did say I'll keep an AGENTS.md or CLAUDE.md. I keep that SUPER high level. The most depth I'll give is info about example projects (That the llm can reach using an MCP) to follow for architectural patterns.
JohnBooty · · focus · HN ↗
Another concrete example would be codebase conventions. We prefer lean models. Cross-model concerns go into service objects. Without this direction LLMs tend to default to stuffing too much code into the models themselves, as is the de facto standard for most MVC apps and therefore this is how LLMs are trained.
Even if they could discover our large codebase's conventions flawlessly on their own in every session, this absolutely would require nontrivial work repeated in every session: quite a few turns grepping, conversing with LSPs, or whatever.
So yes... we could describe our conventions in every single session by typing it right into the prompts... right after typing the directions to find the test credentials... and the other ten or twelve things we'd be telling it every single time...
kaydub · · focus · HN ↗
Even with your documentation, the llm is gonna do a lot of that grepping and discovery. Unless you have such comprehensive documentation that it's basically code itself... in which case, it should just use the code as documentation.
The good news is that I don't think what you guys are doing is going to be really bad. It's just not near optimal and it's creating this weird dogma around all these "tools"
JohnBooty · · focus · HN ↗
kaydub · · focus · HN ↗
Locally? Our devs can do whatever they want. For anything I'm working on, for local deployments and testing, I generally build it out so it's set in Makefile/Taskfile. If you're running your determinstic tests locally it should really just be a single command that bootstraps everything and runs the tests, or updates a certain part of your app and runs the tests. The LLM shouldn't need credentials in that instance.
Regardless I'm deploying a local stack of some sort all the credentials and everything are probably stored locally or in an env var where the llm will have access. So if I was having the LLM drive a browser during active development, it can look at the files.
Sure, I guess we could put in the top level markdown more details about this... but why? It takes little time or context for it to figure it out. We have shit change so frequently that it's just something else we have to maintain.
soltanov · · focus · HN ↗
kaydub · · focus · HN ↗
You guys are just reinventing the wheel and it's fucking square.
johnyzee · · focus · HN ↗
nullsanity · · focus · HN ↗
[dead]
MaurizioFratell · · focus · HN ↗
Genuine question: how is this different or better than existing context and info-about-the-project management systems like for example GSD (formerly get shit done) and others?
dyauspitr · · focus · HN ↗
nelaggy · · focus · HN ↗
ozgung · · focus · HN ↗
There is a sprint based roadmap doc. We do sprint planning and each sprint has multiple sessions with session numbers. There is a backlog document to keep track of all the bugs and feature ideas. There are many ADR documents (Architectural Decision Records) which I admit can become stale in time, but Claude is quick to notice because it actually reads these. They are the main contract between us. For a new feature an ADR is written before writing any code or an old one is amended.
I assumed this was a common workflow because I just prompted Claude to setup the project using its own best practices for a lightweight document-driven development.
I believe asking the model itself is better than using a random repo. These things have access to and trained on an enormous number of projects. They know what they need.
jwpapi · · focus · HN ↗
I only do runbooks for things that I have to run again in the future and potentially mentioned (could be better in code, too tbh, but somehow I like to store these as runbooks, as sometimes infrastucture changes are mixed with a lot of explanation, scraping operations or something)
One off scripts are done in /tmp/
mk_chan · · focus · HN ↗
Documentation was always a proxy for the code so the people who didn’t have context could get started. Obviously if reading the code and reading the doc have the same cost, the doc is completely worthless. This was always the case. You could never depend only on the docs. It’s always the code that actually mattered.
_zoltan_ · · focus · HN ↗
tahaazizi · · focus · HN ↗
[dead]
aarontian0 · · focus · HN ↗
singh_abinashi · · focus · HN ↗
[dead]
trinsic2 · · focus · HN ↗
nialv7 · · focus · HN ↗
IndiaInfraNotes · · focus · HN ↗
[dead]
syngrog66 · · focus · HN ↗
marcus_holmes · · focus · HN ↗
Inside the docs folder is an archive folder. Stale and completed plans and docs get moved there.
I generate a schema.md that describes the database schema in LLM-friendly terms that lives in docs.
If there's a gnarly design or implementation design made that the LLM keeps tripping up on, I get it to document that and put it in docs.
The handoff.md doc lives in docs, and gets moved to archive once it's no longer relevant. Same for todo.md
My agents.md instructs the LLM to read the docs but ignore the archive, and acknowledge the docs it has read. So the start of every session gets a "I have read <date>design.md, <date>plan.md, schema.md, notes.md, handoff.md". It's a decent chunk of context but as the OP says, this is necessary.
I also have a skill called "assessing project drift" that looks at the archive and assesses the project on how far the code has moved from the intention, if there are any chunks of code that were relevant but have been superceded and can be removed, and if there's any egregious tech debt that should be dealt with now. The docs archive is really useful for this.
kaydub · · focus · HN ↗
I think a lot of people are stuck on patterns that "used to work" (and I use that loosely because I really feel like LLMs are getting work done IN SPITE of using all these additional skills/plugins/etc). Go ahead, delete all that shit and try building without. Yeah, sometimes you may have to hand hold the LLM, but that hand holding is a LOT less than doing whatever this is.
This all just sounds like a rube goldberg machine you've made.
marcus_holmes · · focus · HN ↗
And yes, very much a Rube Goldberg machine ;)