It's funny, i push back on pull requests because there is too much description now - a 20 line change has pages and pages of generated description, rationalisation for why it is safe, defense of each design decision, analysis of risks and side effects. People are indignant, you're rejecting my change because there is too much documentation? And my response is, I don't have time to read it and you put me in the position where I can't afford not to - because approving the PR implies I did and accepted it. The investment to read all that for the value of a code change that I'm one prompt away from doing myself if I cared is just not high enough. So it's rejected.
You might be surprised how many organisations/ teams now don’t even review pull requests. Claude does coding, Codex do code review, and if developer did 5 round of this before pull request, we could very well just let CI merge.
I'm very curious how this goes long term. I guess we will find out.
My instinct says that these systems will expand their complexity to fully fit the cognitive budget of the agents that coded them and then atrophy the same way human-built systems do at lower cognitive budget. Only this time, because of the larger up front budget, the complexity ceiling will be higher, and the potential depth of the problem may be much much larger. It may mostly manifest as increasing cost over time - the agents grind for longer and longer, iterating over and over to fix all the failing tests, and the breaking point will be where it never converges and you come back to millions of dollars in budget spent and still tests are failing and effective gridlock on system changes.
But this may be all my human-biased fantasy that justifies still taking a role in software development.
> ... these systems will expand their complexity to fully fit the cognitive budget of the agents that coded them and then atrophy the same way human-built systems do at lower cognitive budget. Only this time, because of the larger up front budget, the complexity ceiling will be higher ...
That is exactly how one could describe the relationship between heroin and the users health systems. It’s okay for a little while, until it isn’t but there is no going back. When the risks start telling, lawyers will have a field day. Assisted by LLM’s obviously, which then puts the burden on judges, who then need LLM assistance.
This thing called attention economics has captured my attention quite hard. Digital opiates are everywhere since 10 to 20 years. And in my view, the part that they implicitly of explicitly strive to capture your attention is new.
Well articulated. I tend to think that the outcome will vary hugely by project. We already observe this with some reporting certain models behave in a certain way, whilst others see the opposite.
To use the wooley term “quality”, the top 20% might stand a good chance of making huge strides. But the remaining 80% of projects (in particular the bottom 20%) will atrophy extremely quickly. Yet, these will be the project that many push LLM’s too as their domain/technology is complex and/or outdated. Digital transformations that can be done quickly will be tantalising but ultimately unsatisfactory long term (as you describe).
and this is such a liability imo. I recently got fired for what I think was apparently merging a change to Kubeconfig that both the CTO and DevOps person approved. I made as damn sure as I could that the proposed change was good, but I needed their eyes on it and they clearly weren't, possibly be cause every other repo had a review bot and no special automatic deployment automations. The DevOps person literally commented and said "It's good to merge", who to trust
Others are just implicitly doing this but pretending review still exists.
Everyone is fatigued by endless code review which you get no credit for and has become massively more of a burden.
All PRs are superficially fine now. There are no typos, there is unit test coverage, but there are deeper issues that require massive amounts of effort and time to spot.
Unfortunately, a lot of previous PR reviews already were just gatekeeping, or "presenteeism". People would leave comments about class names or method names, like they couldn't understand what an `apply` method, the only method, meant on a class with that was already named appropriately and did one thing only (to give an oop example). No, the method needed to be renamed `PriceChecks.applyPriceChecksWithTimeConstraints`.
Lots of review comments about various conditions that wouldn't feasibly happen (same shit with claude now).
But then I'd see these same reviewers approving PRs where the bigger design was just fundamentally broken. Oh, we're adding a blocking call on our hot path, but at least the method name makes it very clear that it is blocking.
In general I agree that the current AI reviews are creating too much noise and it is masking these bigger design issues.
IMO, PRs were never the right process to use in tightly-collaborating teams such as most companies. PRs got popular because orgs started using GitHub, and GitHub had made that workflow to fit the needs of open source (and modeled it after what open source was already doing in the days when patches got sent around via mailing lists).
I'm old enough to remember using CVS and then subversion in companies. People would commit straight to main (which was then called "trunk"), because making feature branches and merging them was cumbersome. And, on regular intervals, the person responsible for some corner of the codebase would do a show-and-tell presenting it to peers, but without the sharply defined boundaries of what the code looked like before vs. after some recent set of changes. People might remember some things from the previous show and tell or from first hand experience with that code, but that kind of memory is necessarily fuzzy, and diffs weren't an artefact that was typical to look at. So, these reviews didn't block people, and any comments that came from reviews defined a direction that things should go from here on out. If a corner of the codebase was deemed to be in a bad shape, the blame around that was equally fuzzy.
zmmmmm · · focus · HN ↗
ramshanker · · focus · HN ↗
fantasizr · · focus · HN ↗
classified · · focus · HN ↗
zmmmmm · · focus · HN ↗
My instinct says that these systems will expand their complexity to fully fit the cognitive budget of the agents that coded them and then atrophy the same way human-built systems do at lower cognitive budget. Only this time, because of the larger up front budget, the complexity ceiling will be higher, and the potential depth of the problem may be much much larger. It may mostly manifest as increasing cost over time - the agents grind for longer and longer, iterating over and over to fix all the failing tests, and the breaking point will be where it never converges and you come back to millions of dollars in budget spent and still tests are failing and effective gridlock on system changes.
But this may be all my human-biased fantasy that justifies still taking a role in software development.
rapidfl · · focus · HN ↗
wow this is a beautiful way to put it
wjnc · · focus · HN ↗
This thing called attention economics has captured my attention quite hard. Digital opiates are everywhere since 10 to 20 years. And in my view, the part that they implicitly of explicitly strive to capture your attention is new.
the_gipsy · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
joshghent · · focus · HN ↗
To use the wooley term “quality”, the top 20% might stand a good chance of making huge strides. But the remaining 80% of projects (in particular the bottom 20%) will atrophy extremely quickly. Yet, these will be the project that many push LLM’s too as their domain/technology is complex and/or outdated. Digital transformations that can be done quickly will be tantalising but ultimately unsatisfactory long term (as you describe).
intended · · focus · HN ↗
How many are you seeing / estimating?
brailsafe · · focus · HN ↗
veltas · · focus · HN ↗
brailsafe · · focus · HN ↗
Gigachad · · focus · HN ↗
Everyone is fatigued by endless code review which you get no credit for and has become massively more of a burden.
All PRs are superficially fine now. There are no typos, there is unit test coverage, but there are deeper issues that require massive amounts of effort and time to spot.
munksbeer · · focus · HN ↗
Lots of review comments about various conditions that wouldn't feasibly happen (same shit with claude now).
But then I'd see these same reviewers approving PRs where the bigger design was just fundamentally broken. Oh, we're adding a blocking call on our hot path, but at least the method name makes it very clear that it is blocking.
In general I agree that the current AI reviews are creating too much noise and it is masking these bigger design issues.
datsci_est_2015 · · focus · HN ↗
stakhanov · · focus · HN ↗
I'm old enough to remember using CVS and then subversion in companies. People would commit straight to main (which was then called "trunk"), because making feature branches and merging them was cumbersome. And, on regular intervals, the person responsible for some corner of the codebase would do a show-and-tell presenting it to peers, but without the sharply defined boundaries of what the code looked like before vs. after some recent set of changes. People might remember some things from the previous show and tell or from first hand experience with that code, but that kind of memory is necessarily fuzzy, and diffs weren't an artefact that was typical to look at. So, these reviews didn't block people, and any comments that came from reviews defined a direction that things should go from here on out. If a corner of the codebase was deemed to be in a bad shape, the blame around that was equally fuzzy.