‹ BackHN Continuity

Thread

Agents don't need memory, they need documentation

364 points · 260 comments · kmeh

  1. kaydub · · focus · HN ↗
    You don't need documentation or the 3rd party memory systems. The code IS the documentation.

    All this stuff is LLM rube goldberg machines. It just pollutes context.

    I barely use AGENTS.md/CLAUDE.md these days. And where they remain, it's super basic high level stuff.

    I'm honestly still kicking myself in the ass on many projects where I did something similar to this. I kept tons of markdown docs and decision docs. Now those things are just causing problems because they got stale. Even after having sessions of reconciling documentation, the LLM just gets confused.

    1. le-mark · · focus · HN ↗
      It’s a weird thing isn’t it, the urge to save these artifacts? The worst is when the llm refers to the decision and design in code comments. In my opinion there’s one thing that is worth documenting; tricky architecture or implementation details that are some how counterintuitive to what would have normally been done. But again this can be documented in the code and tests.
      1. kaydub · · focus · HN ↗
        Commit messages are also great for recording the "why" or other details that don't end up in code.
    2. AppleBananaPie · · focus · HN ↗
      I went through the same cycle as well.

      I think it's going to be an incredibly common, maybe universal cycle people will go through working with AI until they realize it doesn't work long term.

    3. smokel · · focus · HN ↗
      Code typically documents the "what" and "how", not the "why".

      Why something exists, and how it connects to the outside world may be documented in comments, but more often than not it isn't.

      1. haukebri · · focus · HN ↗

        [dead]

      2. arcanemachiner · · focus · HN ↗
        And agents are pretty bad at inferring when to do, and often fall back to verbose clutterin the comments.
      3. kaydub · · focus · HN ↗
        The "why" should most often be self evident. If it's not, that's what commit messages are for. Not more markdown files and code comments.
        1. smokel · · focus · HN ↗
          Unfortunately, that is not how many software development projects work. A customer may request certain things, or the software might be part of a larger system.

          Consider working on software for a coffee machine. Why is pin 42 (GRIND) activated every now and then? Would you really want to document that in Git commit messages?

          1. kaydub · · focus · HN ↗
            Yeah, I'm going too far the other way. There are definitely legitimate reasons to comment in code and have documentation. I just feel like we're seeing a MASSIVE amount of documentation now and it's so much that it's mostly worthless.

            I'm seeing decision files that are so big the LLM can't fit it all in context on some projects. Then the LLM makes decisions that revert previous ones and later sessions don't pick that up so it sticks to the original decision. Now in some sessions, every so often I have to remind the LLM, "no, we changed that later, we do it this way now"

            And I'm seeing our knowledge base grow to a completely useless giant mess of stale, outdated, duplicated, or superfluous info. LLMs often pull unrelated info or confuse similar but different things or get old documentation for something that's been updated to new documentation in a different part of the knowledge base. And these are LLMs generating the docs. And we have LLMs and agents reconciling. But it doesn't seem to always get everything.

            For code comments, it's terrible because the comments are starting to get larger than the code. A large chunk of the comment can be discerned from the code itself. Then the comment has details on why that maybe don't quite make much sense. It's like the LLMs start using words in a specific context that doesn't really apply to the word in normal spoken english. Then it will also often include a specific JIRA ticket id, you check the JIRA ticket, you see that yeah, the code was changed because of that JIRA ticket, but it's not really related to the ticket itself, it was just a blocker. But now the comment forever links it to THAT ticket (And then now sometimes the LLM pulls in that ticket with the atlassian mcp).

    4. haukebri · · focus · HN ↗

      [dead]

    5. rectang · · focus · HN ↗
      Haha, all the software devs who hate writing documentation are naturally finding their preexisting beliefs reinforced when the LLM is able to discern intent without docs. An LLM can be spooky impressive at reading minimized or obfuscated code, for example.

      But this article argues that LLMs do better when the context is smaller — when it can understand the totality of the task with as little context as possible. And so having correct API-level docs is greatly advantageous. Anecdotally, this rings true to me — when the local context is good and clear, the LLM writes code matching my intent even when my prompt is sloppy and poorly specified.

      Rejoice! The LLM will write the docs for you, relieving you of most of the work.

      However without intervention, it will do too much and record absurdly verbose docs (similar to how an LLM will relentlessly refactor your code until you instruct it to move in minimal, incremental changesets). You will still need to edit down what the LLM generates.

      1. klodolph · · focus · HN ↗
        100% agree. Most of the things that make development better for humans also make development better for agents… and I think docs are even more important with agents, because of some kind of multiplicative effect. The agents are coding faster, and the benefits of documentation are somewhat more pronounced because of the speed.
        1. trinsic2 · · focus · HN ↗
          Wow its utterly impressive that people dont understand this. I create docs and maps to instruct my LLM and im always keeping my docs up to date..
          1. kaydub · · focus · HN ↗
            What are you measuring to prove this?

            The docs end up stale or you waste a ton of time and tokens keeping them up to date. And you can't even trust the LLMs to keep the docs up to date, you will still need to review it yourself keep a bunch of shit out of them.

            1. joquarky · · focus · HN ↗
              I don't understand how specifying my preferences for how things are done can go "stale".
              1. kaydub · · focus · HN ↗
                "docs and maps" are hardly specifying preferences. I DID say I still often keep a high level AGENTS.md or CLAUDE.md. Where I have them, they're super high-level. In places where we have well defined and followed standards that top level markdown file can be as simple as "check out this other repo for patterns" and where I don't, the markdown file is high level overview of what the app is and then a high level overview of architecture or some preferences. Always a shorter/smaller doc.

                I don't mind SOME documentation. I'm growing extremely frustrated with the absolute MOUNTAIN of docs, comments, decision documents, etc being created. I'm so fucking tired of the LLM rube goldberg machines being built.

        2. kaydub · · focus · HN ↗
          You guys show me a measurable difference on something with docs vs without docs and I'll believe it.

          The docs just make you feel good. They're not worth anything. They just create more work if anything because now you don't only have to fight entropy in your codebase but also your docs.

          1. rectang · · focus · HN ↗
            Uh, you aren't providing numbers either, just appending the same assertion to the end of each branch of the discussion. The result is the same as when the LLM does that: geometric increase in what we have to fight through to reason about the problem.

            With the LLM, though, I have some influence: I can instruct it to be more succinct. I don't have a specific top-level instruction about that at present in any global prefs file, but I've occasionally asked it to summarize and archive and it's done an okayish job of pruning.

            1. kaydub · · focus · HN ↗
              Yes, you're right, I'm not providing numbers. But that's because I'm just using the tool as designed... the vanilla version. I'm using the baseline. I believe anthropic and openai publish plenty for my side of the argument already.

              I'm not the one bolting things on claiming it increases productivity. Why would I ADD stuff without any proof or evidence that it works? I'm simply NOT adding these things.

              1. frankacter · · focus · HN ↗
                >But that's because I'm just using the tool as designed

                While arguing against using the tool as designed.

                <a href="https:&#x2F;&#x2F;claude.com&#x2F;blog&#x2F;using-claude-md-files" rel="nofollow">https:&#x2F;&#x2F;claude.com&#x2F;blog&#x2F;using-claude-md-files

                &quot;CLAUDE.md files solve this by giving Claude persistent context about your project. A well-configured CLAUDE.md transforms how Claude works with your specific project. The file serves multiple purposes: providing architectural context, establishing workflows, and connecting Claude to your development tools. Each addition should solve a real problem you have encountered, not theoretical concerns about what Claude might need.&quot;

      2. JohnBooty · · focus · HN ↗
        Yeah. Basically IMO&#x2F;IME the current best practice is to start the LLMs out with minimal&#x2F;no skills&#x2F;instructions&#x2F;docs.

        Notice the things they struggle with and the small problems (typically, environment issues IME) they repeatedly encounter and re-solve across multiple sessions. That is what your instructions should cover. When possible, move those instructions into skills, so they get loaded into context selectively instead of on every session. (Example: instructions for running specs, placed into a skill that only gets loaded into context when it’s time to run specs)

        Again, this can be automated by the LLMs themselves: both Codex and Claude (and I’m assuming other major harnesses) know how to read their own transcripts and are good at looking for repeated friction and making concrete suggestions to reduce that friction in the future.

        It takes a bit of a time investment on the user’s part, and every now and then you probably should throw it all out and start fresh so that the new batch of instructions can be appropriate for the current state of the repo and the capabilities of whatever model(s) you’re using.

        1. kaydub · · focus · HN ↗
          Honestly, don&#x27;t agree. I think even this is a waste of time and context.

          The LLMs will always have these suggestions on things that will &quot;reduce friction in the future&quot; but I really think the LLM is a sycophant glazing you. Because even when you add those skills or prompts at some point the LLM ignores it or does whatever it wants to do anyways. Better to leave most of that out and deal with it when it arises.

          Even your example, instructions for running specs... none of the current frontier models need this at all.

          1. joquarky · · focus · HN ↗
            Are you using Astra&#x2F;Fable for everything? Not everyone can afford that.
            1. kaydub · · focus · HN ↗
              Nope. Opus and Sol for the most part. Haven&#x27;t even used Astra or Fable.

              At work it&#x27;s all Anthropic. I spent a while on Sonnet models but have been pretty consistent about using Opus now. For personal projects I bounce between Sol&#x2F;Terra and Gemini Flash.

              Most of the models are good enough now. You don&#x27;t need all these documents, memory, plugins, skills, etc. MCPs are good for hooking up to external sources of data (at work we have gitlab mcp, datadog mcp, and our knowledge base mcp... which the knowledge base is mostly worthless now because of all this AI generated documentation, but I digress). I DO still use beads on a lot of projects, but not always. My prompts these days are lazy as fuck: &quot;review this repo, check out this other repo with architectural patterns we should follow and libraries&#x2F;terraform&#x2F;etc we should use, &lt;basic desc of goal&gt;. What do you think? Let&#x27;s review everything and discuss before we start building&quot;

              Oh no, I might have to stop the llm agent on occasion, or review what they did and tell them to fix a couple things. Way better use of tokens and my time than creating some rube goldberg machine. Way better than keeping a giant decision file that goes stale and causes context corruption because the llm only saw the OLD decision and not the UPDATED decision.

            2. JohnBooty · · focus · HN ↗
              I think they&#x27;re using Skynet. They seem to be claiming that part of the reason for not giving the LLM directions is because the LLM won&#x27;t follow them anyway.

              That must be a remarkable model. Smart enough to flawlessly discover everything on its own in every session... but also it just straight up just doesn&#x27;t obey direction.

              Hmmm.

              1. kaydub · · focus · HN ↗
                Wow, what a strawman.

                I didn&#x27;t say give no direction.

                I&#x27;m saying a lot of devs&#x2F;engineers little rain dances aren&#x27;t really bringing the rain.

                Keep a basic CLAUDE.md&#x2F;AGENTS.md that keeps high level details. Have a way to reference other projects&#x2F;apps&#x2F;codebases to use as reference architecture.

                Don&#x27;t keep tons of markdown files. Don&#x27;t keep &quot;decision&quot; documents (these are probably the worst). Don&#x27;t make skills for every little thing (superpowers are dead these days, stop using them). Oh yeah, don&#x27;t have the LLM generate docs for reference later... if it can generate the docs it doesn&#x27;t really need them.

          2. JohnBooty · · focus · HN ↗
            Charitably, I think we must be working on much different kinds of projects.
            1. kaydub · · focus · HN ↗
              Maybe but I doubt it. We&#x27;re all working on pretty similar things, that&#x27;s why the LLMs work so well.

              You&#x27;re just stuck on your dogma.

              1. JohnBooty · · focus · HN ↗
                I understand what you&#x27;re railing against in general: engineers who make these big memory&#x2F;doc systems that quickly go stale and are at best useless and at worst actively harmful. Even more insidiously, these engineers are blind to this fact because maybe some of this stuff was effective 6-12 months ago and they haven&#x27;t reassessed their workflows. Now that&#x27;s dogma. We really do agree on that.

                However. At least in the projects I&#x27;ve worked on, it&#x27;s trivial to read the transcripts and notice that without enough direction, there are certain things the agents will burn lots of tokens to rediscover in every session. That&#x27;s absolutely not dogma.

                1. kaydub · · focus · HN ↗
                  I think there&#x27;s a balance between the tokens burned up rediscovering things vs tokens being used keeping docs&#x2F;decisions&#x2F;skills&#x2F;etc.

                  And I think a lot of places are burning way more on the latter than the former. I think the former is more cost effective due to caching as well. So I HEAVILY lean towards the former especially with how effective I think it is these days.

      3. kaydub · · focus · HN ↗
        Don&#x27;t agree. LLM generated docs are some of the worst because like I said, the LLM never trims, only amends. So we used to do X but now we&#x27;ve had an architectural change or some type of change where we should never do X. Instead of just removing the instructions to do X in the docs, it amends them, &quot;we made the decision to no longer do X because of Y&quot;. Now it has X multiple times in context instead of just not having X in context at all.
        1. icedchai · · focus · HN ↗
          I hesitate to make blanket statements, since this is quickly evolving, but in general I agree. LLM generated docs will quickly degrade to overly verbose, unreadable crap.
          1. kaydub · · focus · HN ↗
            Not just that they get stale, but if you&#x27;re using the LLM to generate the doc you don&#x27;t need it.

            I don&#x27;t think I&#x27;ve ever seen an agent review a doc and then NOT also go look at the code. So skip the middleman, just have the agent look at the code.

          2. joquarky · · focus · HN ↗
            Created a skill to prune them appropriately and run it every evening.
            1. kaydub · · focus · HN ↗
              Nope. Doesn&#x27;t work. Not consistently. Not without overhead or maintenance.
            2. NBJack · · focus · HN ↗
              That&#x27;s a lot of rolling the dice to risk some interesting (and unseen) hallucinations.
    6. ramesh31 · · focus · HN ↗
      Yup. Examples examples examples. All of the descriptive stuff is just nonsense that confuses the point. Makes perfect sense when you remember that these things are not intelligent, but truly just autocomplete on steroids.
    7. enraged_camel · · focus · HN ↗
      &gt;&gt; You don&#x27;t need documentation or the 3rd party memory systems. The code IS the documentation.

      We have heard this nonsense from the &quot;we don&#x27;t need to write comments, code should be self-documenting&quot; types for decades. It was wrong in that context, and it is wrong in this one.

      Code tells you how a system works. It does not tell you why it works that way. That is what memory is for. It exists so that your AI does not keep undoing past decisions when it writes or refactors code.

      1. kaydub · · focus · HN ↗
        I&#x27;ve had the LLMs undo past decisions MORE from the docs than from having no docs.

        Old decisions always end up in the docs. If the LLM gets a whiff of an old decision, but doesn&#x27;t get the update to that decision, well now you&#x27;re doing things the old way again.

    8. dregitsky · · focus · HN ↗
      &gt; The code IS the documentation

      I liked this advice when humans wrote code. Though even then I&#x27;d urge people to write meaningful commit messages that capture the &quot;why&quot; of what they did, so no one tramples their intent by mistake.

      But not sure it works in an age where most code is LLM-generated. Especially if that code is not even reviewed by humans (irresponsible or not, it&#x27;s happening), and commit messages are also generated by AI. I think something is needed to separate &quot;what did the human operator intend&quot; from what the agent went and built.

      I do agree that this gets way overengineered. My approach has been more or less what you stopped doing though - committing all our timestamped &quot;plan&#x2F;implementation docs&quot; and &quot;investigation docs&quot; that document what the user wanted + empirical findings, and making all prior session transcripts searchable. It&#x27;s seemed mostly helpful? For whatever reason I haven&#x27;t run into many staleness problems so far.

      1. kaydub · · focus · HN ↗
        Commit messages are GREAT.

        I&#x27;m mainly aiming my frustrations at all the markdown files being committed, all the additions to knowledge bases, all the comments in the code (especially the ones referencing specific JIRA tickets). This stuff isn&#x27;t helpful, it gets hella stale. I&#x27;ve had the LLM fuck up plenty due to these docs and comments.

        Commit message ARE EXACTLY where architectural decisions or nuance should go. Not another fucking .md or more comments.

    9. jjfoooo4 · · focus · HN ↗
      What about big projects, where much of the code is not written yet?
      1. kaydub · · focus · HN ↗
        I did say a high-level CLAUDE&#x2F;AGENTS.md is fine.

        Make a high level .md, let it rip, iterate. You can give specifics in your original prompt.

        I don&#x27;t need to document the framework or the libraries etc. Pre-code, I tell it in the prompt one time. After it inits the project it&#x27;s in the code.

    10. mmcnl · · focus · HN ↗
      I don&#x27;t understand. Code doesn&#x27;t capture the context in which decisions were taken: why is code the way it is? What is important? What is not? How can agents make correct decisions without knowing context that cannot be inferred from code?
      1. kaydub · · focus · HN ↗
        Most of the &quot;why&quot; should be self-evident.

        If it&#x27;s not, that&#x27;s what commit messages are for.

    11. hosh · · focus · HN ↗
      Code as documentation works better when the code is declarative or a DSL. These capture intent and promises (as in promise theory) better.

      When it is not, it has to be reasoned out and does not work well for documentation.

      Other things that code and tests alone do not capture well:

      - promises (as in Promise Theory) made to other parties. Claude already has PT in its training data.

      - Constraints-inducing-properites, as in Roy Fielding &#x2F; Christopher Alexander. While tests, and property testing can capture properties, there is no formal connection to the constraints that induces fhem. By constraints, I am not talking about business requirements and business value — those are better understood through Promise Theory. I am talking about things like at-least-once delivery or total ordering (from append-only constraint). Claude already has Fielding’s dissertation and Alexander’s works and ideas in its training data.

      - grammars, as in pattern panguages (not just patterns) a la Alexander &#x2F; Fielding are also not captured in code alone. These tell both humans ans AI how to extend a pattern, and how to identify anti-patterns (when they violate a constraint-inducing-property)

      - LLMs are trained with many different worldviews and bounded contexts at the same time, and is very capable of translating across it. However, these need to be soelled out, otherwise it would talk in whatever it infers

      Specifications written for the exact way components are wited together run into that stale doc problem. Although it takes much more human attention and token burn to describe things in terms of pattern language and promise theory, it becomes easier over time. The actual implementation plan tends to fall out more cleanly when all those other stuff are at least considered. This is where I have been spending most of my time.

    12. chaostheory · · focus · HN ↗
      Code can get much larger than the documentation that summarizes it. It also doesn&#x27;t cover intent or rationale. Even if you inexplicable don&#x27;t want documentation, at the very least use something like gitnexus to map out your code because relying on code alone isn&#x27;t good enough
      1. kaydub · · focus · HN ↗
        That just sounds like a code-smell to me.
        1. chaostheory · · focus · HN ↗
          [delayed]
    13. 01100011 · · focus · HN ↗
      It&#x27;s frequent for SWEs to make blanket statements with considering the vast space of issues other people face that they don&#x27;t have experience with or awareness of. Anyway, I&#x27;m not going to tell you what you do or don&#x27;t need, only what worked and didn&#x27;t for me.

      In my codebase it is difficult to get agreement on comments and documentation so rather than rely on it I adapted. One of the first things I did when I succumbed to agentic development was to point codex at the code and ask it to generate a high level description of where important files, such as our public API, reside, what the hierarchy is, what the code does, etc. In my case, this level of documentation is fairly static if I avoid implementation details. So now I have a handful of agent files in my tree and it seems to save quite a few tokens and improve my results. I frequently have other devs ask me how I get such good results when doing agentic reviews of their changes(always my first step now before I start my human review). I also include instructions in the agents files instructing the agent to maintain the agent files if any relevant changes are made. It seems to work quite well for me.

      1. kaydub · · focus · HN ↗
        This is one of the things I REALLY don&#x27;t get.

        If you got the LLM to generate the docs, they don&#x27;t need the docs.

    14. mgfist · · focus · HN ↗
      Idk I just can&#x27;t agree. Code is the what but it doesn&#x27;t tell you the why. There are so many times where at first glance the code seems suboptimal or bad or wrong, and it&#x27;s only when you learn of some constraint somewhere else that it begins to make sense.

      All code is written under constraints, and most constraints live outside the code.

      1. kaydub · · focus · HN ↗
        Most often the &quot;why&quot; should be self-evident.

        When it&#x27;s not, there are commit messages.

        Please for the love of god, quit generating markdown files (especially having the LLM generate the file, because if it could generate it, it doesn&#x27;t need it), quit generating &quot;decision&quot; docs, and stop having more comments than code.

    15. mike-akdeniz · · focus · HN ↗

      [dead]

    16. ChimpWithHat · · focus · HN ↗
      I partially agree, way too many people are cargo culting overly complex AI workflows with little empirical data. My framing is a bit different though, I consider code the spec and actually keep a decent amount of docs for higher level concepts. So far this is working well for me across Claud and Codex.
      1. kaydub · · focus · HN ↗
        High level docs are fine and can be okay or I can see them as being helpful. I did say I&#x27;ll keep an AGENTS.md&#x2F;CLAUDE.md. There also are SOME docs and sometimes the LLM makes the docs anyways. I&#x27;m not an absolutist, but I&#x27;m REALLY fucking tired of seeing these giant .md files, &quot;decision&quot; documents, more comments than code, etc.

        It&#x27;s a LOT of cargo-culting overly complex AI workflows. It&#x27;s devs&#x2F;engineers making rube goldberg machines.

        Everyone is doing their own little rain dance and when it rains they say &quot;see, I told you it works&quot;

    17. shinokami · · focus · HN ↗

      [dead]

    18. JohnBooty · · focus · HN ↗
      In general, as LLMs get smarter, “less is more” becomes increasingly true. You’re polluting their context with dozens or hundreds of instructions, all of which the LLM tries to satisfy. However,

          You don’t need documentation […or…] memory
      
      …oh heck no. Easy to miss at Claude and Codex’s default detail level but if you read your actual session transcripts, you are almost certain to notice the LLM solving lots and lots of the same little problems over and over again.

          The code IS the documentation
      
      First, lots of things cannot be learned from the code.

      Trivial example: LLMs struggled with QA on our app. They didn’t know how to find the seeded test accounts. They would create new ones and do it wrong, or find the seeded test accounts but not know their passwords because they were encrypted, so they’d change the passwords but not tell the other agents. Shitloads of tokens burned. It was a no-brainer to just give them the credentials in a skill that gets loaded when they do QA.

      You could say that’s an environment issue, not a code issue. But as far as actual code IME at a minimum we have to tell the agents our general repo structure and architecture patterns.

      The days of “you are the world’s greatest Python programmer, write good code” or whatever are over (if that style of prompting even worked in the first place) and as I said less is more. But, still….

      1. kaydub · · focus · HN ↗
        &gt; You could say that’s an environment issue, not a code issue

        Yeah, that&#x27;s exactly what I&#x27;d say. Or maybe you&#x27;re just approaching the problem now.

        You have to prompt these QA agents, correct? Why not give the instructions on where to get credentials in the original prompt?

        &gt; But as far as actual code IME at a minimum we have to tell the agents our general repo structure and architecture patterns.

        I did say I&#x27;ll keep an AGENTS.md or CLAUDE.md. I keep that SUPER high level. The most depth I&#x27;ll give is info about example projects (That the llm can reach using an MCP) to follow for architectural patterns.

        1. JohnBooty · · focus · HN ↗

              Why not give the instructions on where to get 
              credentials in the original prompt?
          
          We don&#x27;t push code unless it&#x27;s been QA&#x27;d, and the LLMs need to know the credentials every time they do automated QA. It&#x27;s very rare for me have a session in that repo where they wouldn&#x27;t need that info. So it&#x27;s a great candidate for AGENTS.md

          Another concrete example would be codebase conventions. We prefer lean models. Cross-model concerns go into service objects. Without this direction LLMs tend to default to stuffing too much code into the models themselves, as is the de facto standard for most MVC apps and therefore this is how LLMs are trained.

          Even if they could discover our large codebase&#x27;s conventions flawlessly on their own in every session, this absolutely would require nontrivial work repeated in every session: quite a few turns grepping, conversing with LSPs, or whatever.

          So yes... we could describe our conventions in every single session by typing it right into the prompts... right after typing the directions to find the test credentials... and the other ten or twelve things we&#x27;d be telling it every single time...

          1. kaydub · · focus · HN ↗
            I&#x27;m assuming these QA agent are automated, probably in CI&#x2F;CD. And I&#x27;m assuming the prompt that runs these QA agents is from a codebase in VC. Why would I put QA instructions into the codebase&#x27;s repo as a top level markdown file and not in the pertinent part of my QA process?

            Even with your documentation, the llm is gonna do a lot of that grepping and discovery. Unless you have such comprehensive documentation that it&#x27;s basically code itself... in which case, it should just use the code as documentation.

            The good news is that I don&#x27;t think what you guys are doing is going to be really bad. It&#x27;s just not near optimal and it&#x27;s creating this weird dogma around all these &quot;tools&quot;

            1. JohnBooty · · focus · HN ↗
              Do you not run your tests locally? Regardless of what&#x27;s happening or not happening in CI&#x2F;CD, I would think that most workflows involve running the tests locally. We do TDD more or less and for the browser-based tests, the LLMs need to know how to log into stuff.
              1. kaydub · · focus · HN ↗
                No, these kind of tests aren&#x27;t generally run locally. Not to get to prod at least. These are all deterministic tests to get to prod.

                Locally? Our devs can do whatever they want. For anything I&#x27;m working on, for local deployments and testing, I generally build it out so it&#x27;s set in Makefile&#x2F;Taskfile. If you&#x27;re running your determinstic tests locally it should really just be a single command that bootstraps everything and runs the tests, or updates a certain part of your app and runs the tests. The LLM shouldn&#x27;t need credentials in that instance.

                Regardless I&#x27;m deploying a local stack of some sort all the credentials and everything are probably stored locally or in an env var where the llm will have access. So if I was having the LLM drive a browser during active development, it can look at the files.

                Sure, I guess we could put in the top level markdown more details about this... but why? It takes little time or context for it to figure it out. We have shit change so frequently that it&#x27;s just something else we have to maintain.

    19. soltanov · · focus · HN ↗
      Versioned documentation is inspectable, reviewable, and easier to correct than opaque recalled snippets. Documentation should describe current truth; append-only events can preserve history.
      1. kaydub · · focus · HN ↗
        This is what commit messages are for.

        You guys are just reinventing the wheel and it&#x27;s fucking square.

        1. soltanov · · focus · HN ↗

          [dead]

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.