‹ BackHN Continuity

Thread

How to Write with an LLM

769 points · 420 comments · joeriddles

  1. semiquaver · · focus · HN ↗
    This would sound insane to me from two years ago but I have recently started insisting on writing all my own commit messages and pull request descriptions. I do usually have an agent review them for factual accuracy, but not rephrase them.

    It slows things down a bit, but in the best possible way. It has helped immensely to improve the depth of my understanding of the agent-generated code. When agents are doing everything its way too easy to “skim” diffs and not really absorb them.

    I always prided myself on my technical writing, and commit messages and PRs were a great place to hone that skill. I found that I missed it and my work is better now I’ve reclaimed that part of my old job back.

    1. BeetleB · · focus · HN ↗
      At work, I simply don't allow LLMs to make commits. :-)
    2. hannasanarion · · focus · HN ↗
      This currently is my team's only AI-use policy and I thik it's working out fantastically.

      "No AI PR Descriptions" is a great rule because it does 3 things:

      1. It's a hard binary that's easy to recognize and enforce

      2. It establishes personal ownership for the submitter, making it psychologically difficult to submit something you don't understand. If a person has to describe what a change does, how, and why, then they must have looked at it and made a real attempt to understand it, because you can't map a territory you have never seen. This alone prevents "claude-code run amok" scenarios where devs have completely abdicated responsibility that are a real pain to clean up.

      3. It informs the reviewer of the change's intent. A human PR tells reviewers what a thing is intended to do, as well as the authors belief about what it does, which they can use as a framework to critique the actual substance. It becomes easier to notice missed edge cases, code behavior that subtly differs from intent, or changes that would block off or complicate an intended future project direction or reduce maintainability.

      And as a bonus, at time of writing, AI pr descriptions basically always suck. They are usually full of irrelevant implementation details, misplaced emphasis, weird presuppositions that seem to come out of nowhere and are really annoying to read, etc. That won't be true forever, they are always getting better, but it is true now, so that's a temporary 4th benefit.

      1. vova_hn2 · · focus · HN ↗
        Great comment, no idea why it was flagged
      2. dawnerd · · focus · HN ↗
        Our rules are: short title that should match the task in the ticket tracker, link to the task in the body as the first item, pr summary that’s human readable, screenshots of any front end changes at all the breakpoints. The screenshot requirement is huge, it’s caught a lot of people skipping visual testing locally. Even if they have an llm generate the screenshot we know it’s been at least tested by something and allows for a quick check before someone even wastes their time looking at code.
    3. LeafItAlone · · focus · HN ↗
      A big win of LLMs in my book is the absolute reduction of commit messages with just “fixes” or “updates”. Commit messages have become more meaningful and useful, even if far from perfect.

      We have one dev who uses LLMs to write the code, but still commits by hand. Most of his messages are of the type above, and none of them are useful.

      1. semiquaver · · focus · HN ↗
        Yeah, I know people who have the same zero-explanation `git commit -mfix` style that they did pre-AI and am baffled. If you can’t be arsed to explain yourself, let the agent do it. It literally takes less work to tell the agent to commit for you and they will always do better than -mfix. I don’t understand it either other than obstreperous “become ungovernable” attitude.
      2. this_user · · focus · HN ↗
        I'm not sure that Claude's "Realigned the shape of the load-bearing ownership gate to reduce the blast radius of the design contract; confirmed, not assumed" is more meaningful than "fix".
        1. 0x696C6961 · · focus · HN ↗
          The long-ass Claude commit messages & PR descriptions suck. But they are 100% better than "fix".
          1. ericbarrett · · focus · HN ↗
            Mildly disagree. "Fix" is exasperating but immediately tells me I need to look at the diff. With unconstrained Claude spew, I need to wade through three levels of deep fried LLM-speak before realizing...I need to look at the diff
            1. atif089 · · focus · HN ↗

              [dead]

            2. atif089 · · focus · HN ↗
              This.
            3. bitwize · · focus · HN ↗
              "Deep fried" is the perfect analogy. LLMs have been trained on themselves so many times that their output is the linguistic equivalent of many rounds of JPEG compression.
            4. danielbln · · focus · HN ↗
              I don't get this thread. Just tell the agent to use conventional commit messages and to keep it nice and tight. Are you all just raw dogging agent output with no alignment/conventions?
              1. ericbarrett · · focus · HN ↗
                Sadly yes, a lot of people are
              2. rudedogg · · focus · HN ↗
                > Just tell the agent to use conventional commit messages and to keep it nice and tight.

                I have this in my claude.md along with guidance on (not) writing comments but it’s still dumps multi-paragraph comments of Claude speak everywhere.

                I think a lot of us are talking past eachother, but what I and I think others are complaining about is doing things well, and being minimal. LLMs don’t really do that yet, they can’t look at and simplify a codebase well, even with guidance. They always seem to add rather than subtract. And eventually it becomes an issue. I think they were actually better in this regard with like Opus 4.8, and are now getting even worse (but score better, and are more autonomous).

                At home I use Codex, and it’s better as far as language goes at least.

                Anyway, it’s like the logic piece necessary to do good work is still missing, and hurt by recent reinforcement fine-tuning. And it leads to massive bloat since more comments = higher scores, but walls of text and the tokens (or cognition if you’re the sad human reading them) aren’t free.

          2. burnished · · focus · HN ↗
            How? One looks like someone cared about it and the other can be dismissed at a glance
        2. winrid · · focus · HN ↗
          "fix stupid shit" is all you need :)
        3. skissane · · focus · HN ↗
          > I'm not sure that Claude's "Realigned the shape of the load-bearing ownership gate to reduce the blast radius of the design contract; confirmed, not assumed" is more meaningful than "fix".

          For PR/commit descriptions, I mainly use Claude Sonnet 4.5. It isn’t perfect, but it produces significantly less of this weird gibberish than 5.x models or even Opus 4.x do

          I also use an iterative process in which it writes the description, I read it, and then either manually edit it or ask it to make changes

          1. my-next-account · · focus · HN ↗
            I use Astra at Very High, and shit is still bad. It doesn't actually understand anything, so it often says things which are clearly not needed to be stated. Recently, I've learned that I have very high standards for these things. For example, "fixes" as a commit msg just is NOT acceptable and would never fly where I work.
            1. manmal · · focus · HN ↗
              Astra is such a mixed bag. It makes some amazing reviews and sometimes architecture suggestions that I like. But it’s also lazy and will just make up things.
              1. senderista · · focus · HN ↗
                hence adversarial review
                1. manmal · · focus · HN ↗
                  I’m familiar with how those are used, but not sure what you mean in this context.
                  1. adastra22 · · focus · HN ↗
                    Have another model (or even another instance of the same model) review the output of the first.

                    Models will hallucinate. They are also quite good at spotting hallucinations in other models' output (with some more hallucinations thrown in). With a threshold for confirmation, and a few iteration loops, you arrive at a fixed point where every claim is supported.

            2. rrr_oh_man · · focus · HN ↗
              Very high does not improve the model, fyi.
              1. my-next-account · · focus · HN ↗
                Wat, what am I paying for then?
                1. mitxela · · focus · HN ↗
                  Dario's yacht and FOMO.
                2. rrr_oh_man · · focus · HN ↗
                  Arguably, your LLM provider might be paying you, in a sense
                  1. oblio · · focus · HN ↗
                    As OpenAI's IPO troubles seem to indicate, that's an OpenAI skill issue.
                    1. oblio · · focus · HN ↗
                      Touchy OpenAI employees? :-)))
                3. adastra22 · · focus · HN ↗
                  A few things, but generally more chain of thought before generating a response. So the model is tuned to think more. Given that what it outputs for this task is a summary of its thinking, tuning it to think more will just make a more verbose, less useful commit message.

                  Tune your model parameters to what is right for the task, not the highest you can afford.

            3. jester997 · · focus · HN ↗
              Honestly you probably want a model that has only been trained on language and literature. Nothing from online discourse.

              And even then… writing is personal expression. Here people are talking about commit messages. That’s fine but AI doing writing for anyone and I WILL NOT READ IT unless it’s literally basic tech manual.

              We read to hear and engage with people’s thoughts. If someone outsources that to AI then they should be shunned.

        4. ben_w · · focus · HN ↗
          Could go either way, TBH.

          So long as there are in fact those things, and so long as it didn't sneak something else in there at the same time, it being just on the knife-edge between sense and word-salad is better than "fix".

          Buuuuuut far to often it says it fixed a bug I reported, when it only touched one superficial failure mode rather than the root cause.

          Yesterday's issue: Why is zoom/pan randomly failing? It told me it was because it was applying a transformation matrix with every input and sometimes JavaScript gave it a non-invertible matrix (why?) which then propagated NaNs everywhere and you can't update a matrix filled with non-numbers.

          Why was it doing that in the first place? Seems to be because it's too motivated to perform quick wins and not sufficiently motivated to do good engineering.

          Good thing this was just a game editor. Spiky intelligence: superhuman on some dimensions, total noob on others.

        5. TeMPOraL · · focus · HN ↗
          It is, because it includes the key words such as "ownership gate" and perhaps others, which makes it infinitely more informative than just "fix", by pointing at what was in scope, and what wasn't.
        6. jasonlotito · · focus · HN ↗
          I've never had Claude spew that out in a commit message. That being said, I also use my tools properly. Your comment sounds like the type of thing someone who puts "don't make any mistakes" at the end of a prompt and gets upset when it doesn't come out perfect. Like, someone adding "make it better than AAA" is going to do something. At this point, it's really just telling on yourself that this is the output that you get.
        7. [deleted] · · focus · HN ↗

          [deleted]

        8. [deleted] · · focus · HN ↗

          [deleted]

        9. raincole · · focus · HN ↗
          FYI I never had LLM writing such a commit message. So as they say, skill issue.
      3. aprilthird2021 · · focus · HN ↗
        You've gotta be joking. Every AI commit message, summary, and test plan is full of extraneous garbage I didn't need to read (and a lot of it is not understandable)
        1. LeafItAlone · · focus · HN ↗
          Fix your harness. Provide rules, format, and good examples.

          You probably had those for the humans in your team for when new hurts join, right? Just direct your LLMs to them and you’ll have a much better experience.

          Just like you can’t expect a new junior dev to know what to do without direction, the same applies with LLMs.

          1. skydhash · · focus · HN ↗
            The issue is with the people sending code for reviews, not people reading the slop commit message.
          2. aprilthird2021 · · focus · HN ↗
            I can do that, but other people I work with won't, but I have to review or block their shit
        2. theultdev · · focus · HN ↗
          It's not for you, it's for the sessions later. Personally I let it save every last tidbit in the commit details, why not.
          1. aprilthird2021 · · focus · HN ↗
            It's for me if it's sent to me for review!!
      4. randito · · focus · HN ↗
        Disagree. I've seen so many PR slop grenades -- they are essentially unreadable. I'd prefer a paragraph or two human written explanations. I suspect people are just taking the default AI text -- which is way too wordy.
      5. Terr_ · · focus · HN ↗
        Sometimes minimal information is better than extrapolated information, especially when the extrapolation mentally painful to read and has decent probability of being wrong with inaccuracies that waste people's time.

        Also:

        1. It's often a premature extrapolation from diffs and tickets which you could probably generate later, at need, and likely with even better results.

        2. If those commit messages can ever influence future work, then it's just carcinogenic cargo-cult cruft. Pruning outputs (future inputs) is important to prevent weird unwanted drift.

      6. FooBarWidget · · focus · HN ↗
        I can relate to that. Most human commit messages aren't good because the author can't be bothered to write a good one, or doesn't know what a good message is supposed to be, or is bad at writing.

        But AI commit messages are still bad. Way too verbose and focused on the wrong level: that of code mechanics. That's just wrong. Messages should focus on the level of intent and design, with the primary purpose being to aid human review. They should include a high level overview of the change, decisions, caveats, information not obvious from reading the diff.

        So I wrote a skill that captures these principles and allows the agent to even research past related commits and to ask focused questions in order to uncover the intent rather than guessing and writing a bad one.

        Now the AI writes better commit messages than it used to, and even writes better than most humans (who can't be bothered to write a good one). Not better than a good manual message, but you can't have everything.

        But sometimes even the good writers are tired or didn't think things through. Being able to compare with the AI's version is still useful.

        Here is the skill for anyone interested: <a href="https:&#x2F;&#x2F;github.com&#x2F;FooBarWidget&#x2F;ai-skills-and-principles&#x2F;blob&#x2F;main&#x2F;skills&#x2F;commit-and-pr-messages&#x2F;SKILL.md" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;FooBarWidget&#x2F;ai-skills-and-principles&#x2F;blo...

        Depends on my &quot;documentation principles&quot; skill: <a href="https:&#x2F;&#x2F;github.com&#x2F;FooBarWidget&#x2F;ai-skills-and-principles&#x2F;blob&#x2F;main&#x2F;skills&#x2F;documentation-principles&#x2F;SKILL.md" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;FooBarWidget&#x2F;ai-skills-and-principles&#x2F;blo...

      7. hobo123 · · focus · HN ↗
        But that&#x27;s a code review issue. Nobody should get away with shoddy commit messages like that.
        1. vova_hn2 · · focus · HN ↗
          I don&#x27;t think that many teams review commit messages as a part of code review. At least, I&#x27;ve never seen it &quot;in the wild&quot;.
          1. hobo123 · · focus · HN ↗
            True, is mostly about the code being merged, but when reviewing a PR they might suggest that the dev write better commit messages in the future.
            1. dawnerd · · focus · HN ↗
              We have standards we ask devs to follow but I’m not going to make someone go back and change a commit message after the fact. I more care about pull request body copy and I think the rules we have for that have been pretty well adhered too. Don’t even need ai to enforce it.
        2. semiquaver · · focus · HN ↗
          Not all workplaces are high functioning in the way you imagine. Until recently a shocking amount of software development did not use source control.
      8. bob1029 · · focus · HN ↗
        I am way more interested in the intentions of the author than anything that can be interpolated from the code or token space.
        1. vova_hn2 · · focus · HN ↗
          Exactly! If I want a summary of the code, I can easily generate it myself.
        2. LeafItAlone · · focus · HN ↗
          LLMs can provide the intentions of the code they write. To say they can’t is saying that they can’t fix bugs, which they very obviously can do effectively. So, just tell them to write the intention in the way you want it. If you are implicitly providing the direction in the prompt, make that explicit - you then it will have the context for the change while working and for the commit message and you’ll likely get a higher quality output.
      9. atmosx · · focus · HN ↗
        An LLM can read the diff and do the summary on the fly. What it cannot do, is explain the reason behind a change. That was always that valuable part of the commit msg. The rest can deducted, ofc it&#x27;s much easier to read a short list or a descriptive title.
        1. [deleted] · · focus · HN ↗

          [deleted]

      10. vova_hn2 · · focus · HN ↗
        I think that the idea of generating a commit message based on the content (diff) of the commit is fundamentally wrong.

        Even before coding harnesses become mainstream, a lot of tools offered to automatically generate commit messages based on the diff (I think JetBrains IDEs started to offer it very early) and I always cringed when I&#x27;ve seen it.

        The reason why I don&#x27;t like diff-base commit messages is because they are redundant. If I want an LLM-generated summary of the diff, I can easily generate it myself, there is absolutely no reason to put it in the commit message.

        What I would like to see in the commit message is some additional context that is not a part of the diff. I don&#x27;t need to read what has changed, because I can already can see it in the diff (or get an LLM to summarize it for me). But I often do need to understand why this change was made. What were they trying to achieve? That&#x27;s the important part that is not contained in the diff itself. And this is the part that diff-base commits rarely contain.

        I think that even something simple like a link to a Jira (or other bugtracker ticket) is much more helpful than diff summary. Maybe instead of &quot;fix&quot; or &quot;update&quot; you could write just a couple of words about what are you trying to fix and why does it need updated. Still better, than a diff-summary.

        1. miki123211 · · focus · HN ↗
          This is why, if you have the agent generate the commit message at all, it has to be the same agent that wrote the code, in the same session.

          I usually have it lead with a paragraph or two of context on why the change was made (though I often use it as just a tool to turn my rambling on that question into proper English), followed by a couple paragraphs of &quot;abstract&quot; explaining the change (because an LLM-generated summary reviewed by the person who made the commit is better than something which the person reading can get themselves).

          1. adastra22 · · focus · HN ↗
            I don&#x27;t see how that would improve the situation? If I do it with the same agent&#x2F;same context, it tends to just report what it did. Which is the diff, repeated in prose.
            1. miki123211 · · focus · HN ↗
              The agent is usually told why it was making the change (I at least tend to start my prompts with a description of the problem we&#x27;re solving), so, if prompted correctly, it can lead with &quot;there was a bug, we root-caused it and here&#x27;s what the root-cause was&quot;, then explain what the fix entailed.
        2. [deleted] · · focus · HN ↗

          [deleted]

        3. grayxu · · focus · HN ↗
          true
        4. pc86 · · focus · HN ↗
          &gt; But I often do need to understand why this change was made.

          This is the reason (and IMO the only reason) code should be commented. The why can often be hundreds of characters explaining technical limitations at the time, business considerations, or even just &quot;we know this will break in a non-trivial number of instances but it&#x27;s an incremental step forward to our end goal so we&#x27;re just doing it.&quot;

      11. zmmmmm · · focus · HN ↗
        Commit messages as you describe (&quot;fixes&quot;, &quot;updates&quot;) are inappropriate in any professional context and some coaching should occur to the people doing them.

        I have the opposite issue - some of my team members now submit mini-essays generated by the LLM. Like 300-500 word commit messages with everything from the essence of the change up to philosophical design trade off discussions.

        Like most writing, what is left out is as important as what is included.

    4. initsecret · · focus · HN ↗
      same. also—from the reviewer end—LLM generated PR descriptions are loooooooooooooong.
    5. dyauspitr · · focus · HN ↗
      Why do you want to understand the agent generated code though? That would make you the bottleneck. Shouldn’t you just concentrate on checking for optimal outcomes, exhaustively trying to find&#x2F;prompt edge cases and profiling performance?
      1. codechicago277 · · focus · HN ↗
        It’s difficult if not impossible to know where to look for edge cases or performance problems without understanding the code.
      2. semiquaver · · focus · HN ↗
        Letting agents run wild with a codebase and no humans understanding it is a recipe for disaster across so many axes.

        You are not a serious person if you recommend that.

        Check back in a couple years and I’m sure they’ll be there but today’s frontier models 100% absolutely are not ready to fully own a nontrivial production codebase with no human involvement.

      3. intended · · focus · HN ↗
        Process vs outcomes.

        If your work has little liability then you can afford to not care beyond “does it work”.

        If you have to worry about quality and ensuring you don’t get sued, you make sure the process works.

        If it has to maintainable, you you need the mental model to be present in someone’s head.

      4. tiborsaas · · focus · HN ↗
        I do check most of agent code, not because I really care about every tiny detail but to catch if it&#x27;s going to shoot itself (and my project) in the foot. If I only check the result and it&#x27;s OK, it still doesn&#x27;t mean that in two features I won&#x27;t totally blow up the app.

        I&#x27;m a happy little bottleneck, what&#x27;s wrong with that? How do you even know how to ask it to build stuff for you if you don&#x27;t understand what foundations are you building on?

        Writing commit message (aka. what has changed) gives me a sense of control that I know what&#x27;s happening and I can confidently build the next thing.

      5. FuckButtons · · focus · HN ↗
        Oh no, the bottleneck is understanding what the fuck you’re doing! Think of the time you could save!
    6. cgriswald · · focus · HN ↗
      Using LLMs is like managing people. It’s difficult in some ways and easier in others. You should trust them to a degree but also know what is going on and if you’re “trusting” them out of laziness you’re doing it wrong.
    7. jdkoeck · · focus · HN ↗
      Two years ago? I don’t get it, coding agents have been compelling for barely a year.
      1. holoduke · · focus · HN ↗
        I would say from march this year.
      2. adastra22 · · focus · HN ↗
        Cursor and aider were definitely compelling two years ago. Nothing compared to now, of course, but still better than raw dogging it.
    8. solarkraft · · focus · HN ↗
      I still manually edit almost every word for utmost precision in anything I expect a human to read since LLM prose is hard to read and often subtly wrong or misleadingly worded (those false contrasts ...).

      Commit messages are not exactly in this category for me, I view them mostly as a work log to be later inspected by another LLM to gather context. I do usually review them before approving a given plan, so I do care about their structure and content, but find the LLM sufficiently competent at writing them.

      1. Terr_ · · focus · HN ↗
        For me, the use-case of commit messages is to help a human discover when some kind of thing might have changed, as distinct from the commits before and after it. Importantly, the human is doing a subjective scan for something that sounds relevant.

        If they already knew what function or detail is involved, they would be already be doing a &quot;all commits that touched this line&quot; filter, and my comment about affecting $thing would probably be superfluous.

        I also don&#x27;t need to tell them such details in the commit, because that is better expressed by the actual diff.

        So in a sense, I&#x27;m trying to provide good keywords, about transactions, errors, logging, button color, whatever.

    9. maximg68 · · focus · HN ↗
      Makes sense. I started asking Claude to grill me on the PR to make sure I understand it.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.