‹ BackHN Continuity

Thread

Agents don't need memory, they need documentation

363 points · 259 comments · kmeh

  1. CapitalistCartr · · focus · HN ↗
    This is something I've been fooling with a lot lately. Reading his solution, it look to me like his objections apply to his own solution. The sharpest critique he makes of RAG is that agents can't search for what they don't know. A markdown "brain" has the same problem. How does the "agent" know which documents are relevant before it starts? The index files retrieval done by the agent instead of by embeddings doesn't escape the problem. The same goes for staleness (which for me seems like a constant chase). He criticizes memory systems for treating the past as truth, but documents go stale too (a lot). The fix of having the agent update what's outdated, is the same job he's ridiculing the dreamers and background daemons for doing. He says "putting it to the test"; where's the test? He says "only five of the many problems"; if there's so many, show me, don't just say it. He's absolutely right about auditability, but for me at least Claude uses a regular markdown (MD) file I can read just fine. So every memory plugin on the market does not work the same way.

    This is a first draft; his github is better than his article. Looking through it, Consult actually works. The agent doesn't pick documents blind. Every scope has a catalog file that describes each document: what it covers, when to open it. These catalogs seem to load in to the start of each session, so the agent gets a little map without reading every file. Code navigation seems the same. Each index document has a short description and a "read_if", and subindexes are opened when their condition matches the job. This looks pretty well laid out, which I would never have guessed from the article.

    1. crazygringo · · focus · HN ↗
      > A markdown "brain" has the same problem. How does the "agent" know which documents are relevant before it starts?

      I don't know, but I have a pretty standard (I think?) setup, and Claude manages to find every relevant file every time. But I've also only used Claude for greenfield projects, where "documentation is primary, and code flows from documentation" is the philosophy.

      I have CLAUDE.md describe all the types of documentation files and the directory structure. And then Claude is pretty aggressive (automatically) about always inserting cross-references everywhere. So a feature description will reference the ADR's that it implements, the ADR's say what feature implements them. A code file will make reference to the "implementation design" document that describes the motivation behind which iOS elements were chosen, how the animation is defined in a particular way that doesn't break another animation, and so forth. So I've really never run into a situation where Claude failed to read a file it should have. I've been pleasantly surprised.

      I would say that the one really big thing I've had to learn is to teach Claude both in CLAUDE.md and in the header of every top-level design document, that keeping documentation current and in sync is paramount. Because its default seems to be to keep history and append, e.g. by default it will take a section of a document and mark it "[DEPRECATED]" and add the new version below. So my instructions are pretty clear in having it always be aggressive in maintaining current state only, always replace rather than append. And if there's anything we want to save from the previous approach (e.g. we did it X way previously and it failed because Y), then just add that as a new short note in the new current-state text, possibly with a pointer to a commit or tag or something.

      So this seems to solve both recall and staleness in my projects at least.

      The only thing I still haven't found a solution for is numbering. Claude is always giving everything numbers, like F23 for feature 23. But I'm always changing the order of things, inserting new things, deleting things, so I wind up with a sequence of development work that goes in order like "Phase 9", "Phase 9b", "Phase 9e", "Phase 11", "Phase 12". I'm halfway ready to abandon numbers entirely and just start giving things names from noun collections instead, so every feature is named after an animal, every ADR is named after a kitchen implement, or something. Or just four-digit hex codes chosen at random. Curious if anyone else has found what works.

      1. hedgehog · · focus · HN ↗
        Numbered tasks are fine so long as the number doesn't determine the completion order. I use a task tree system that has some rules about target task size (essentially 150k tokens or 45 minutes) and then lets the agent manage adding, splitting, dependencies, priority, etc. Seems to work fine up to around 1000 tasks.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.