‹ BackHN Continuity

Thread

Agents don't need memory, they need documentation

365 points · 261 comments · kmeh

  1. kaydub · · focus · HN ↗
    You don't need documentation or the 3rd party memory systems. The code IS the documentation.

    All this stuff is LLM rube goldberg machines. It just pollutes context.

    I barely use AGENTS.md/CLAUDE.md these days. And where they remain, it's super basic high level stuff.

    I'm honestly still kicking myself in the ass on many projects where I did something similar to this. I kept tons of markdown docs and decision docs. Now those things are just causing problems because they got stale. Even after having sessions of reconciling documentation, the LLM just gets confused.

    1. rectang · · focus · HN ↗
      Haha, all the software devs who hate writing documentation are naturally finding their preexisting beliefs reinforced when the LLM is able to discern intent without docs. An LLM can be spooky impressive at reading minimized or obfuscated code, for example.

      But this article argues that LLMs do better when the context is smaller — when it can understand the totality of the task with as little context as possible. And so having correct API-level docs is greatly advantageous. Anecdotally, this rings true to me — when the local context is good and clear, the LLM writes code matching my intent even when my prompt is sloppy and poorly specified.

      Rejoice! The LLM will write the docs for you, relieving you of most of the work.

      However without intervention, it will do too much and record absurdly verbose docs (similar to how an LLM will relentlessly refactor your code until you instruct it to move in minimal, incremental changesets). You will still need to edit down what the LLM generates.

      1. JohnBooty · · focus · HN ↗
        Yeah. Basically IMO/IME the current best practice is to start the LLMs out with minimal/no skills/instructions/docs.

        Notice the things they struggle with and the small problems (typically, environment issues IME) they repeatedly encounter and re-solve across multiple sessions. That is what your instructions should cover. When possible, move those instructions into skills, so they get loaded into context selectively instead of on every session. (Example: instructions for running specs, placed into a skill that only gets loaded into context when it’s time to run specs)

        Again, this can be automated by the LLMs themselves: both Codex and Claude (and I’m assuming other major harnesses) know how to read their own transcripts and are good at looking for repeated friction and making concrete suggestions to reduce that friction in the future.

        It takes a bit of a time investment on the user’s part, and every now and then you probably should throw it all out and start fresh so that the new batch of instructions can be appropriate for the current state of the repo and the capabilities of whatever model(s) you’re using.

        1. kaydub · · focus · HN ↗
          Honestly, don't agree. I think even this is a waste of time and context.

          The LLMs will always have these suggestions on things that will "reduce friction in the future" but I really think the LLM is a sycophant glazing you. Because even when you add those skills or prompts at some point the LLM ignores it or does whatever it wants to do anyways. Better to leave most of that out and deal with it when it arises.

          Even your example, instructions for running specs... none of the current frontier models need this at all.

          1. JohnBooty · · focus · HN ↗
            Charitably, I think we must be working on much different kinds of projects.
            1. kaydub · · focus · HN ↗
              Maybe but I doubt it. We're all working on pretty similar things, that's why the LLMs work so well.

              You're just stuck on your dogma.

              1. JohnBooty · · focus · HN ↗
                I understand what you're railing against in general: engineers who make these big memory/doc systems that quickly go stale and are at best useless and at worst actively harmful. Even more insidiously, these engineers are blind to this fact because maybe some of this stuff was effective 6-12 months ago and they haven't reassessed their workflows. Now that's dogma. We really do agree on that.

                However. At least in the projects I've worked on, it's trivial to read the transcripts and notice that without enough direction, there are certain things the agents will burn lots of tokens to rediscover in every session. That's absolutely not dogma.

                1. kaydub · · focus · HN ↗
                  I think there's a balance between the tokens burned up rediscovering things vs tokens being used keeping docs/decisions/skills/etc.

                  And I think a lot of places are burning way more on the latter than the former. I think the former is more cost effective due to caching as well. So I HEAVILY lean towards the former especially with how effective I think it is these days.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.