‹ BackHN Continuity

Thread

Agents don't need memory, they need documentation

364 points · 260 comments · kmeh

  1. CapitalistCartr · · focus · HN ↗
    This is something I've been fooling with a lot lately. Reading his solution, it look to me like his objections apply to his own solution. The sharpest critique he makes of RAG is that agents can't search for what they don't know. A markdown "brain" has the same problem. How does the "agent" know which documents are relevant before it starts? The index files retrieval done by the agent instead of by embeddings doesn't escape the problem. The same goes for staleness (which for me seems like a constant chase). He criticizes memory systems for treating the past as truth, but documents go stale too (a lot). The fix of having the agent update what's outdated, is the same job he's ridiculing the dreamers and background daemons for doing. He says "putting it to the test"; where's the test? He says "only five of the many problems"; if there's so many, show me, don't just say it. He's absolutely right about auditability, but for me at least Claude uses a regular markdown (MD) file I can read just fine. So every memory plugin on the market does not work the same way.

    This is a first draft; his github is better than his article. Looking through it, Consult actually works. The agent doesn't pick documents blind. Every scope has a catalog file that describes each document: what it covers, when to open it. These catalogs seem to load in to the start of each session, so the agent gets a little map without reading every file. Code navigation seems the same. Each index document has a short description and a "read_if", and subindexes are opened when their condition matches the job. This looks pretty well laid out, which I would never have guessed from the article.

    1. athrowaway3z · · focus · HN ↗
      What I wish more people would be talking about is that RAG should be considered harmful.

      When you have knowledge distributed in markdown files; finding them puts the path/filename into context as well as some indication of document size. (If its on line 1200 or line 20). This is extremely valuable for picking what ought to be focused on next.

      RAG on the other hand creates the hardest challenge for these models. It instead puts 5 ideas with the highest similarity into the context in full.

      Its the difference between having to remember a set of numbers when in a crowd that's talking about stuff, and having to remember them when the crowd is shouting out random numbers. The similarity in the task makes things harder. SoTA models work despite this, but its extra-gambling while you're already gambling.

      1. yesb · · focus · HN ↗
        >RAG on the other hand creates the hardest challenge for these models. It instead puts 5 ideas with the highest similarity into the context in full.

        I think what you're observing is that there is more to information retrieval i.e. "retrieval" in RAG than slapping everything into a vector database and calling it a day. There's no such requirement in RAG to mindlessly load the k nearest neighbors into your context and see what happens. That's a very rudimentary implementation.

        This markdown system I'd argue is RAG as well. You're just doing the retrieval in a way customized for the problem at hand. If you have a precise method of retrieving the most relevant things, obviously use that rather than a similarity metric. If I'm reading correctly, this markdown system is basically a knowledge graph which is not a new idea.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.