This would sound insane to me from two years ago but I have recently started insisting on writing all my own commit messages and pull request descriptions. I do usually have an agent review them for factual accuracy, but not rephrase them.
It slows things down a bit, but in the best possible way. It has helped immensely to improve the depth of my understanding of the agent-generated code. When agents are doing everything its way too easy to “skim” diffs and not really absorb them.
I always prided myself on my technical writing, and commit messages and PRs were a great place to hone that skill. I found that I missed it and my work is better now I’ve reclaimed that part of my old job back.
A big win of LLMs in my book is the absolute reduction of commit messages with just “fixes” or “updates”. Commit messages have become more meaningful and useful, even if far from perfect.
We have one dev who uses LLMs to write the code, but still commits by hand. Most of his messages are of the type above, and none of them are useful.
I'm not sure that Claude's "Realigned the shape of the load-bearing ownership gate to reduce the blast radius of the design contract; confirmed, not assumed" is more meaningful than "fix".
Mildly disagree. "Fix" is exasperating but immediately tells me I need to look at the diff. With unconstrained Claude spew, I need to wade through three levels of deep fried LLM-speak before realizing...I need to look at the diff
"Deep fried" is the perfect analogy. LLMs have been trained on themselves so many times that their output is the linguistic equivalent of many rounds of JPEG compression.
I don't get this thread. Just tell the agent to use conventional commit messages and to keep it nice and tight. Are you all just raw dogging agent output with no alignment/conventions?
> Just tell the agent to use conventional commit messages and to keep it nice and tight.
I have this in my claude.md along with guidance on (not) writing comments but it’s still dumps multi-paragraph comments of Claude speak everywhere.
I think a lot of us are talking past eachother, but what I and I think others are complaining about is doing things well, and being minimal. LLMs don’t really do that yet, they can’t look at and simplify a codebase well, even with guidance. They always seem to add rather than subtract. And eventually it becomes an issue. I think they were actually better in this regard with like Opus 4.8, and are now getting even worse (but score better, and are more autonomous).
At home I use Codex, and it’s better as far as language goes at least.
Anyway, it’s like the logic piece necessary to do good work is still missing, and hurt by recent reinforcement fine-tuning. And it leads to massive bloat since more comments = higher scores, but walls of text and the tokens (or cognition if you’re the sad human reading them) aren’t free.
> I'm not sure that Claude's "Realigned the shape of the load-bearing ownership gate to reduce the blast radius of the design contract; confirmed, not assumed" is more meaningful than "fix".
For PR/commit descriptions, I mainly use Claude Sonnet 4.5. It isn’t perfect, but it produces significantly less of this weird gibberish than 5.x models or even Opus 4.x do
I also use an iterative process in which it writes the description, I read it, and then either manually edit it or ask it to make changes
I use Astra at Very High, and shit is still bad. It doesn't actually understand anything, so it often says things which are clearly not needed to be stated. Recently, I've learned that I have very high standards for these things. For example, "fixes" as a commit msg just is NOT acceptable and would never fly where I work.
Astra is such a mixed bag. It makes some amazing reviews and sometimes architecture suggestions that I like. But it’s also lazy and will just make up things.
Have another model (or even another instance of the same model) review the output of the first.
Models will hallucinate. They are also quite good at spotting hallucinations in other models' output (with some more hallucinations thrown in). With a threshold for confirmation, and a few iteration loops, you arrive at a fixed point where every claim is supported.
A few things, but generally more chain of thought before generating a response. So the model is tuned to think more. Given that what it outputs for this task is a summary of its thinking, tuning it to think more will just make a more verbose, less useful commit message.
Tune your model parameters to what is right for the task, not the highest you can afford.
Honestly you probably want a model that has only been trained on language and literature. Nothing from online discourse.
And even then… writing is personal expression. Here people are talking about commit messages. That’s fine but AI doing writing for anyone and I WILL NOT READ IT unless it’s literally basic tech manual.
We read to hear and engage with people’s thoughts. If someone outsources that to AI then they should be shunned.
So long as there are in fact those things, and so long as it didn't sneak something else in there at the same time, it being just on the knife-edge between sense and word-salad is better than "fix".
Buuuuuut far to often it says it fixed a bug I reported, when it only touched one superficial failure mode rather than the root cause.
Yesterday's issue: Why is zoom/pan randomly failing? It told me it was because it was applying a transformation matrix with every input and sometimes JavaScript gave it a non-invertible matrix (why?) which then propagated NaNs everywhere and you can't update a matrix filled with non-numbers.
Why was it doing that in the first place? Seems to be because it's too motivated to perform quick wins and not sufficiently motivated to do good engineering.
Good thing this was just a game editor. Spiky intelligence: superhuman on some dimensions, total noob on others.
It is, because it includes the key words such as "ownership gate" and perhaps others, which makes it infinitely more informative than just "fix", by pointing at what was in scope, and what wasn't.
I've never had Claude spew that out in a commit message. That being said, I also use my tools properly. Your comment sounds like the type of thing someone who puts "don't make any mistakes" at the end of a prompt and gets upset when it doesn't come out perfect. Like, someone adding "make it better than AAA" is going to do something. At this point, it's really just telling on yourself that this is the output that you get.
semiquaver · · focus · HN ↗
It slows things down a bit, but in the best possible way. It has helped immensely to improve the depth of my understanding of the agent-generated code. When agents are doing everything its way too easy to “skim” diffs and not really absorb them.
I always prided myself on my technical writing, and commit messages and PRs were a great place to hone that skill. I found that I missed it and my work is better now I’ve reclaimed that part of my old job back.
LeafItAlone · · focus · HN ↗
We have one dev who uses LLMs to write the code, but still commits by hand. Most of his messages are of the type above, and none of them are useful.
this_user · · focus · HN ↗
0x696C6961 · · focus · HN ↗
ericbarrett · · focus · HN ↗
atif089 · · focus · HN ↗
[dead]
atif089 · · focus · HN ↗
bitwize · · focus · HN ↗
danielbln · · focus · HN ↗
ericbarrett · · focus · HN ↗
rudedogg · · focus · HN ↗
I have this in my claude.md along with guidance on (not) writing comments but it’s still dumps multi-paragraph comments of Claude speak everywhere.
I think a lot of us are talking past eachother, but what I and I think others are complaining about is doing things well, and being minimal. LLMs don’t really do that yet, they can’t look at and simplify a codebase well, even with guidance. They always seem to add rather than subtract. And eventually it becomes an issue. I think they were actually better in this regard with like Opus 4.8, and are now getting even worse (but score better, and are more autonomous).
At home I use Codex, and it’s better as far as language goes at least.
Anyway, it’s like the logic piece necessary to do good work is still missing, and hurt by recent reinforcement fine-tuning. And it leads to massive bloat since more comments = higher scores, but walls of text and the tokens (or cognition if you’re the sad human reading them) aren’t free.
burnished · · focus · HN ↗
winrid · · focus · HN ↗
skissane · · focus · HN ↗
For PR/commit descriptions, I mainly use Claude Sonnet 4.5. It isn’t perfect, but it produces significantly less of this weird gibberish than 5.x models or even Opus 4.x do
I also use an iterative process in which it writes the description, I read it, and then either manually edit it or ask it to make changes
my-next-account · · focus · HN ↗
manmal · · focus · HN ↗
senderista · · focus · HN ↗
manmal · · focus · HN ↗
adastra22 · · focus · HN ↗
Models will hallucinate. They are also quite good at spotting hallucinations in other models' output (with some more hallucinations thrown in). With a threshold for confirmation, and a few iteration loops, you arrive at a fixed point where every claim is supported.
rrr_oh_man · · focus · HN ↗
my-next-account · · focus · HN ↗
mitxela · · focus · HN ↗
rrr_oh_man · · focus · HN ↗
oblio · · focus · HN ↗
oblio · · focus · HN ↗
adastra22 · · focus · HN ↗
Tune your model parameters to what is right for the task, not the highest you can afford.
jester997 · · focus · HN ↗
And even then… writing is personal expression. Here people are talking about commit messages. That’s fine but AI doing writing for anyone and I WILL NOT READ IT unless it’s literally basic tech manual.
We read to hear and engage with people’s thoughts. If someone outsources that to AI then they should be shunned.
ben_w · · focus · HN ↗
So long as there are in fact those things, and so long as it didn't sneak something else in there at the same time, it being just on the knife-edge between sense and word-salad is better than "fix".
Buuuuuut far to often it says it fixed a bug I reported, when it only touched one superficial failure mode rather than the root cause.
Yesterday's issue: Why is zoom/pan randomly failing? It told me it was because it was applying a transformation matrix with every input and sometimes JavaScript gave it a non-invertible matrix (why?) which then propagated NaNs everywhere and you can't update a matrix filled with non-numbers.
Why was it doing that in the first place? Seems to be because it's too motivated to perform quick wins and not sufficiently motivated to do good engineering.
Good thing this was just a game editor. Spiky intelligence: superhuman on some dimensions, total noob on others.
TeMPOraL · · focus · HN ↗
jasonlotito · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
[deleted] · · focus · HN ↗
[deleted]
raincole · · focus · HN ↗