What I've found is that AI allows lazy and incompetent developers to be more lazy and more incompetent. This then has the effect that product quality suffers more, faster. As a result of the sheer amount of code now being pushed out, code reviews, a thing that previously somewhat prevented lazy and incompetent developers from pushing out horrible code, is effectively dead in the water since no human can actually review such amounts of code realistically anymore. Some companies have adopted AI to review code, which, well ... you have AI make code, AI review code ... I hope you can see the stupidity here if you expect to see any deterministic results at all.
I guess time will tell if the consumer will adapt to the lower quality of products, allowing companies to justify the existence of lazy and incompetent developers, or if the consumer will push back, forcing companies to increase the quality of their developers.
Note: I use AI every day and it is entirely possible to create high quality software with it, so long as you are not lazy and incompetent.
I mean these guys are not even pretending to be reviewing the code.
It just gets “reviewed” by an LLM, which will find a nitpick while ignoring the huge fire in the core of the design, force the planner to make even more sloppy code to cover for an irrelevant test case. Rinse old tokens and repeat until you hit limits.
For example, I recently got brought in to help with quality on a large-scale system that had been ported to a new platform with the help of coding agents. The project was completed and declared operational in record time, but soon after the business discovered that:
1. The promised scalability improvements did not materialize. Instead, it got worse.
2. Observability had been lost. The telemetry was no longer trustworthy.
3. Users stopped trusting it because it was producing incorrect outputs.
What I ended up discovering was that, while it scrupulously kept existing automated tests passing, any behavior that wasn't explicitly covered by a test was free to change any which way. And there were plenty of small things that weren't explicitly covered. Perhaps because the original authors thought they were so obvious and commonsense that they didn't need one, perhaps because mistakes happen. The why doesn't matter. The point is that reality is messy and imperfect, so giving someone a chance to look at things and think, "Huh, that's funny..." is an essential part of defense in depth.
The real worst part was, this whole replatforming was a huge waste of time, anyway. The improvements they were looking for could easily have been accomplished with some controlled incremental changes to the original system. Mostly just removing a few basic and well-known performance antipatterns.
But way back at the outset, the person in charge of the project asked their agent, "What's the best way to X," and the agent gave them a trendslop answer about how Y alternative technology is more scalable and we should just port to that. It was convincing and they were under intense time pressure to just ship some code because leadership is bought into the AI hype and now has the patience of a 4 year old, so they just went with it.
Where's the great software, then? I'm genuinely asking: where is it? Because I can't find it, and it's been close to a year since AI for programming has started to take off.
Somebody posted that GitHub status summary: in ten years, they averaged 10 incidents per month — this includes the last 12 months. In the last six months, average is 22.
i have unlimited tokens and i throw Fable / Astra at everything. They suck ass still for anything nontrivial. I could commit that garbage but if I kept doing it, I will end up with a ball of mud only Fable / Astra can grok..convenient for Dario and SamA..
I've asked Codex with GPT 6 Astra to *review" a one time benchmarking script for any mistakes (built by Claude Code using Opus 5.5) and it refactored the shit out of it claiming all sorts of stuff without even asking about the context in which it was developed.
If I was to employ them to review the code without giving each the same baseline multi-page prompt, they go into endless loop of "improvement" with no end goal in sight.
More and more frequently, I instruct frontier models to stop and go back to the task at hand.
Frontier models are subjectively worse at this, in my experience. Very intelligent but very prone to expanding scope. I need to spend more time prompting to get good results. Maybe I'm just working on boring stuff that doesn't require "frontier" intelligence?
Can the users of the software even keep up at that point? We may have reached diminishing returns on software production, and not enough impact on the rest of the process.
So where is all that new software? My laptop and phone run essentially the same software as 2 or 3 years ago. Yes there were some minor updates to some apps, but nothing faster than in the years prior.
The only real updates I've seen to anything have been AI features... so all this AI is only being used to add AI to stuff. Most of which the average person doesn't seem to want or use.
If I understand it correctly, Firefox has some new-ish features that are marketed as AI-made, like tab grouping and tab comments... which also just get in the way for me.
The problem is that as a lay user of Firefox, I don't know how you could even make it better in terms of features (I could see it being faster etc.).
Firefox has added a lot of anti-features, like constant pop-ups that randomly trying to explain the UI to me. It’s almost as bad as the needless browser onboarding process in Edge (on servers this a huge pain).
There are a few one-shotted personal single page apps, yes. Mostly useless beyond an initial "that's cool" factor. Pre-LLM that was occupied by the "I hacked this in a weekend" niche.
I remember fighting with business about this in a startup way before LLMs, more code, velocity or motion doesn't mean more value.
If I put on my optimist hat for a moment, I hope LLM's would finally make this obvious for people, more stuff !== more value. 250k lines of code a week isn't a flex IMO, it doesn't really mean anything out of context.
So far I mostly see a metric sh.. ton of meta- and meta-meta-projects re-wrapping AI wrapper tools, with sloppy slogans like "One Model, Five Harnesses. Combined." or "You run in the park. Rrrunnnrr.ai runs your AI." Looking inside, out of 250k it's often 200k of verbally incontinent self-explaining comments; or "smart" redesign of builtins.
No new browser, no new iOS clone than runs on Android, no new easy to use DaVinci, no new CAD suite, no $5 SolidWorks clone, no redesigned K8s, no 10x performance speedup in Linux kernel.
That's okay. Reviewing the code will become the agents' job as well.
A couple more step functions in model capability of the type we've seen in the past year, and there will pretty much be no reason for humans to be involved in the development process at all. All humans would need to do is communicate clearly what needs to be made and flag problems as they come up.
I've stopped calling out Claude mistakes on team meetings because this is so true.
I mean, sure, I could have predicted in what ways an LLM would fuck up, but there's just so many ways I can't keep up.
We just had a major production issue because someone's LLM wrote queries against dev databases. Which are very obviously dev databases because they are labelled with dev in the name, and in the table descriptions. AI reviewer didn't catch it, neither did the human reviewer for that matter.
> All humans would need to do is communicate clearly what needs to be made and flag problems as they come up.
Kinda what i'm doing already, but for the young startup I'm at that's surprisingly tons of work. I miss the days we wrote code by hand boy those were fun 8.5 hours workdays.
> a thing that previously somewhat prevented lazy and incompetent developers from pushing out horrible code
Brings to mind this classification
<a href="https://en.wikipedia.org/wiki/Kurt_von_Hammerstein-Equord#Classification_of_officers" rel="nofollow">https://en.wikipedia.org/wiki/Kurt_von_Hammerstein-Equord#Cl...
"""I distinguish four types. There are clever, hardworking, stupid, and lazy officers. Usually two characteristics are combined. Some are clever and hardworking; their place is the General Staff. The next ones are stupid and lazy; they make up 90 percent of every army and are suited to routine duties. Anyone who is both clever and lazy is qualified for the highest leadership duties, because he possesses the mental clarity and strength of nerve necessary for difficult decisions. One must beware of anyone who is both stupid and hardworking; he must not be entrusted with any responsibility because he will always only cause damage"""
The problem here is that AI is consistently one of the four things: hardworking. This makes it very efficient at transforming "stupid and lazy" inputs into "stupid and hardworking" outputs.
Now instead of 90% stupid and lazy (harmless, useful for grunt work) you have 90% stupid and hardworking (aggressively causing damage).
We developed languages that removed GOTO so that developers don't shoot themselves in the foot. We will surely develop harnesses that will ensure that majorly occurring problems are solved before they hit production.
Since the output of human software work is code and AI software work is _also_ code they are both liable to shoot themselves in the foot in the same manner.
You see this already, LLMs are a lot more reliable in statically typed languages with strong memory guarantees (like typescript or rust) than in weaker languages.
IMO the only way LLM code can avoid most of the pitfalls of human code is if we make new programming languages targeted at being used by LLMs exclusively. Think of languages with very strong methods for formal proofing and stuff like that.
The problem is that even if said language was invented, it would still fail catastrophically when integrated with systems not made in said language. We are very lucky that relational databases already provide a somewhat high level of formal proofing in this regard.
Said language would be impossible to parse by humans, kinda like assembly where you can parse what an isolated piece of assembly code is doing, but if you can't comprehend a somewhat large pure-assembly codebase as a whole.
> You see this already, LLMs are a lot more reliable in statically typed languages with strong memory guarantees (like typescript or rust) than in weaker languages.
It’s common advice to wire in deterministic checks to your workflow with LLMs - static languages aren’t inherently better for LLMs, it’s that LLMs produce better code when given deterministic feedback, such as compiler results.
Yes, this was my point, formal proof static analysis makes LLM output better. Therefor a language with far higher requirements on formal proofing (to the point it becomes very difficult for humans to understand) might end up being better to use by an LLM.
Of course it is not that simple as a large part of how effective LLMs are is due to training data which is hard to get for a new language.
I've written very large assembly codebases, it's no different than writing in any other language. You have functions you call with inputs and outputs - though usually those are pointers to memory locations. The program is not one long function, you can split it up into different files and folders and keep everything very well organized and easy to understand and reason about.
And another corollary is the formerly golden lazy and clever are also transformed into lazy and productive because they no longer need to apply their cleverness to get results...
I think with proper oversight you could get a reasonable facsimile of a Clever + Hard-worker from an AI-enhanced Lazy + Clever worker. (Maybe just hopefully thinking of myself.)
Someone who knows that 1 + 1 = 2 will not decide that it's suddenly 3 unless we start accounting for health problems. Making mistakes is not the same as non-deterministic.
> Someone who knows that 1 + 1 = 2 will not decide that it's suddenly 3 unless we start accounting for health problems.
And?
The p(that kind of error) is pretty small now. At what point does a probability coming out of an LLM look like "knowing", such that spitting out the wrong answer despite that probability looks like a health problem, a typo, or even just boredom? (Thinking of the Lizardman constant here: <a href="https://en.wiktionary.org/wiki/Lizardman%27s_Constant" rel="nofollow">https://en.wiktionary.org/wiki/Lizardman%27s_Constant)
It's a continuum for both them and us, even if the mechanism is wildly different.
> Making mistakes is not the same as non-deterministic.
i.e. when the dismissal is "non-deterministic" when it should be "Making mistakes", is itself a mistake.
Lol, humans make such absurd mistakes (and worse) all the time through simple typos, which is effectively random. The key for 2 is right next to the key for 3, after all.
I think that actually reinforces the distinction being made. An LLM’s nondeterminism is in the generation process: given the same prompt and model state, sampling can produce different outputs. That doesn’t mean the underlying fact itself becomes nondeterministic.
A human who knows 1+1=2 can still say “3” because they misread the question, misspoke, were distracted, or made some other cognitive error. Likewise, an LLM can output “3” because the generation process selected an incorrect continuation. Those are both errors in producing an answer, not evidence that 1+1 somehow has multiple answers.
So yes, human mistakes and LLM sampling are mechanistically different. If your argument is that LLMs and humans can both make mistakes, then major question here is why are we building out huge amounts of infrastructure at unsustainable spending levels to enable LLMs to make the same mistakes as humans.
> If your argument is that LLMs and humans can both make mistakes
It's not, I'm just pointing out that LLMs won't make that mistake.
You could ask an LLM what 1+1 is, and the number of times it says "3" is so small that it makes no sense to worry about it. It will phrase the response differently each time; that's the nondeterminism. But it won't say "3".
> then major question here is why are we building out huge amounts of infrastructure at unsustainable spending levels to enable LLMs to make the same mistakes as humans.
Yes, if we ignore everything else, that seems like a reasonable question. But let's not ignore everything else, like the fact that LLMs are much more productive than humans and likely already make fewer mistakes than the average programmer.
>You could ask an LLM what 1+1 is, and the number of times it says "3" is so small that it makes no sense to worry about it...
I think the disturbing fact is that you can take a frontier model with all the intelligence of humanity, and make it say 1 + 1 = 3, but specifically training for it...
A human with that much knowledge will refuse that attempt. There in lies the difference..
Yes, I fail to see anything meaningful. If you move the goalposts and say “I have invented a human that cannot be convinced in any way to give a wrong answer” then what’s the point of that in this discussion, really?
>Hit them with a stick until they answer as you told them to.
Obviously, for this purpose, human should not have any feelings (because LLMs don't have), so can't feel pain. Or else the comparison can't work.
LLMs have something functionally equivalent to pain, in this regard at least.
The weights are updated depending on if the feedback was positive or negative.
It has a functional effect similar to that which pleasure and pain have with us. Not identical, so far as I know there's not been any reports of any machine learning model that is into BDSM, but for the most part functionally similar.
> Someone who knows that 1 + 1 = 2 will not decide that it's suddenly 3 unless we start accounting for health problems.
But this is plainly false. This kind of unforced error occurs all the time.
For example, once when I was in high school I traced an error in my math homework to an intermediate calculation of "2 + 2" as being "3". There was no reason.
What we can say about humans is that, if they know that 1 + 1 = 2, (a) they are unlikely to change their mind about this in any kind of lasting or permanent way, and (b) the rate at which they will mistakenly produce other values for 1 + 1 is very low. But it will happen occasionally, and when it does happen, "they just suddenly decided on the wrong value" is an extremely accurate description of what that looks like.
By not reviewing, reading, or understanding the code generated by agentic LLMs the output is effectively like a compiler. However, a compiler has deterministic behaviour that can be repeated and verified.
The behaviour/output of an LLM is not like that. Ask an LLM to create a dashboard to show games by genre and it will generate different results with each run, and each model/model version produces wildly different results.
I've seen LLMs do something correct 98% of the time then randomly do something crazy that a human would never do because we have continual learning
As humans we don't have our memory reset multiple times per day
> I've seen LLMs do something correct 98% of the time then randomly do something crazy that a human would never do because we have continual learning
I've seen humans vote for Brexit, re-elect Trump, ask questions clearly already answered in an FAQ, try to pull on a door labelled "push", and insist on giving me homeopathic silicon dioxide pills* that cost £5** for a 10-12 gram packet.
Continual learning is a difference, but not by itself a reason to care about "deterministic results".
Nor, indeed, correct results.
> As humans we don't have our memory reset multiple times per day
Humans need sleep well before they can read a million tokens' worth of written text. We're more like 300k tokens if you're actually reading and not skimming for 16 hours straight.
Again, different (in soooo many ways), but this isn't a relevant difference when the topic is "deterministic results".
If I do the same task 100 times I'm not going to suddenly do it crazily different at time 101 because I've built in the memory of how to do it
There is no RNG involved when I decide to push vs pull the unlabeled door to my building every morning, it becomes deterministic because its baked into memory
You can put stuff in context to deal with this but you can't do that for everything, its not practical and you would blow the context window
> If I do the same task 100 times I'm not going to suddenly do it crazily different at time 101 because I've built in the memory of how to do it
That "because" is an unimportant detail.
You, me, everyone, we all fail tasks with some non-zero probability, including tasks we've done many times before. This is literally why typos happen. We don't always catch the typos we make, this is literally why spellcheck exists. We don't always catch logical errors etc., this is literally why professional writers work with professional editors. We don't always spot errors in code we write; this was why compiler errors have to exist at one level, why unit tests and integration tests exist at another, and why despite automated tests we have QA departments.
Statistically speaking, you as a human will have made more comprehension errors reading that paragraph than even LLMs a few years back.
> There is no RNG involved when I decide to push vs pull the unlabeled door to my building every morning, it becomes deterministic because its baked into memory
There's a lot of random inputs into human behaviour. How much they become random outputs is a similar question as for AI.
Your memories are not "baked" at any point; they're made of and by cells doing chemistry. Your entire body is a trillion small bags of chemicals that communicate with each other by leaking out signalling molecules and electrochemical gradients into each other's water supply, and this goes wrong sometimes.
Small (and apparently random) groups of neurons in your (and all animal) brains briefly enter sleep-like states even while you remain apparently awake. This starts well before you even feel tired.
"optimize this code", "fix this code", "extend this code", "add this feature", "find errors and patch them", "find bugs and fix them", "rewrite this from python to rust".
This is all that's needed to actually use LLMs nowadays. How is it a "multiplier" rather than an "equalizer"?
> How is it a "multiplier" rather than an "equalizer"?
Because without the responsible human engineer in the loop, it'll all gradually decay in a cascade of edge-cases. This happens with human written code as well (every "we'll replace this prototype before we ship" you've ever worked on), but with LLMs it happens at 10-100x the rate.
Do you use the word "equalizer" in this context to mean that AI has made the playing field equal for both competent developers and laypeople? Do you reckon that competence plays no role these days?
If that's how you create software then you belong to the lazy and incompetent group in my book. I provide AI with valuable context such as code coverage information, architecture analysis, test requirements, important "gotcha's" that a competent engineer would know about in their architecture or system etc. I'm still very much the person who comes up with the solutions. For me AI is replacing the code editor, it's not replacing the thinking.
The skill floor has definitely been lowered, but if this were actually true then firms would be replacing senior software positions with entry level ones, not the other way around.
In optimizing my game I noticed framerate hitching even after efficient algorithms were in place for expensive stuff, which was caused by shaders not being precompiled consistently or assets not being preloaded in time. The Agent who'd been profiling and optimizing had moved many preloads to a loading screen, which caused a long loading lock, and what it didn't move ahead was loaded and compiled at use, creating slow frames since work was being done on the main game loop.
I instructed the agent to create a speculative pre-warming/compilation priority queue with a per-frame budget, with priority being determined by likelihood signals that the asset or shader will be used soon. Then I had the AI run fully headed games and hunt down causes for frames going over 16ms, and work through them until a batch of games had fewer than 1/1000 frames >16ms and no frames over 60ms after a short initial settling period.
The approach, the metrics, the validation system and the loop were "prompt engineering" above and beyond what I would expect from someone who was merely "vibe coding a game."
> you have AI make code, AI review code ... I hope you can see the stupidity here...
You will be surprised how many times, catches errores made by the AI coding agent. However,as you point, isn't deterministic. And you can guarantee the end results is 100% fine code
> Note: I use AI every day and it is entirely possible to create high quality software with it, so long as you are not lazy and incompetent.
What I in general try to teach the other people about AI: It can be a great tool, but check the results! Especially in the case of engineering: Check and then double check.
> What I've found is that AI allows lazy and incompetent developers to be more lazy and more incompetent. This then has the effect that product quality suffers more, faster.
Yeah. To me it seems very much like the "use dynamic typing for everything" fad. You had a bunch of junior and/or incompetent developers who went around insisting that type declarations are bad, static typing slows down development, you just code so much faster if everything is dynamically typed. And in the context of a new project, they were totally right. It took a few years for the debt to finally catch up, and people realized that these massive, untyped monoliths they had were unmaintainable. Now the two biggest dynamic languages (Python/JavaScript) are effectively typed languages, because nobody uses their untyped variants for serious work.
Dynamic typing still has great uses -- interactive data exploration, putting together quick scripts (though less relevant with AI...), or even just simple prototypes -- but what we tried to do with it at the start, as an industry, was clearly dumb as hell. I suspect we'll look back in 5-10 years and realize that with some of the stuff we're doing with AI, too. It's already happened with things like Gastown.
The web wouldn't have taken off without dynamic typing, PHP first of all (and Python/JavaScript after that). People seem to forget how atrocious it was to write an .asp or .jsp (I think the extension was .jsp) page back in 2003-2005.
I'm seeing this too. I've worked with devs that would previously push PRs that wouldn't work or run correctly. Those PRs wouldn't get merged in. Now they're putting up PRs which seem to work at first glance, but have hidden problems. For example, one guy introduced a huge PR for a visualization and it seemed to work fine, though another dev mentioned to me that we already use recharts and it does 90% of what this guy's PR does (his code does all the drawing logic itself). Maybe AI will get good enough to clean up these kinds of messes, but in the near term I imagine there will be a lot of code bases that will be filling up with dragons.
At a startup I worked, there was an engineer whose code was incoherent and buggy. So, we were literally better off if that engineer did nothing because their net output was negative. Engineers like that become weaponized with LLMs, and negative numbers become larger negative numbers when scaled up.
> Is that the fault of AI or management for not firing them?
How does the system behave in a variety of scenarios including failures and restarts. How is state maintained coherently. There are the kinds of systems problems that an engineer needs to reason through, and if there are bugs in such decisions, they end up becoming costly. I dont expect AI or LLMs to solve these problems at all, since each of them has nuances and tradeoffs which are specific to each system. In short, there is specification complexity in precisely describing system wide behaviors, and unfortunately, there is no lean/tla+ to meaningfully describe systems at scale. You could then ask: How can a system have guaranteed behaviors if they cannot be even stated or proved formally ? The answer to this is how protocols like raft/paxos initially convinced us of their behaviors which is in human review and understanding. That begs the question: How can human review and understanding be reliable, and the answer is that it is not reliable, but humans have ability and processes to continuously learn from experience in the real world. So, our understanding is grounded not only by whats out there in books etc, but also by our own interactions with the world.
Long story short: The responsibility for system-wide behaviors of software systems relies on human review and understanding, which while imperfect can continuously learn.
I've seen similar. They wasted weeks of senior engineering time, between reviews, meetings, and follow up in Slack, only to have the PR closed without merge. The offending individual was eventually moved to another project.
It probably would've been merged if it had a smaller blast radius. It had "fixed" (actually broken) many unrelated tests in the process of making a small update.
> forcing companies to increase the quality of their developers
Just don't. Fire them! AI is better than a thousand devs. What you need is testers that know what to test that AI can't, not code or UX/UI (not talking about playwright here) but business intelligence if that is testable, the things that produce results (profits) and the reason it was asked for in the first place, to solve a problem
If the problem was asked wrongly, the result will be wrong too. Fire devs, then PMs, then IT Managers if they really don't know how to outperform AI, and that's exactly the point, they won't be able to do it in code or tests or reviews, only in intelligence, for now...
I'm with you on this. At my employer, I feel like we are looked down upon if we don't take the lazy approach and let the ai attempt to one shot whatever it is we're working on.
One reason I'm reluctant to hand over all of my work to the ai is I don't want to forget how to program or let my skills deteriorate. Another reason is I don't want to become dependent on ai and find myself in a situation where I'm not able to fly/navigate/land the airplane if my auto-pilot or ai malfunctions or fails.
Then the last reason I don't want to take the lazy approach: When I've done "one shot tests" a lot of times the ai will try and take some lazy half-ass shortcut that we would not accept if it were a human doing the work. A lot of times it just doesn't do what you ask it to do.
Where I've found ai extremely helpful though is asking questions about our codebase, or asking it to build me a function that takes in a, b, c arguments and spits out x, y, z.
AI really is one of the greatest things mankind has ever produced, but I don't think it's so good yet that it can replace humans completely. Using it as a form of leverage though I think is what people should be doing. I suppose we'll see what happens to developers who let the ai take over completely. Some people are arguing that if you don't let the ai takeover completely your career is doomed, but personally I think you might be doomed if you forget how to fly the airplane by hand.
Like I mentioned here elsewhere: AI replaces the IDE/editor, not the thinking. We now just work one abstraction level higher, but your ability to architect solutions is as relevant, if not more so, than ever before. I design test cases, architectural plans, review the output, do verification. AI just writes the code and helps me with research. If your job can be entirely offloaded to AI then I’d question if your job was that needed to begin with, but writing code was never the job, engineering was. These are just tools to do the job with, they aren’t the job itself.
We solve business problems through technology and to me it’s quite concerning how many developers think their job was knowing syntax of a particular language. Nobody besides themselves care about the syntax of a particular language, certainly the business doesn’t care. I question if retaining your knowledge of programming languages is that important anymore, but your ability to read code if needed (after all, most programming languages are quite similar so it’s not that hard), do architectural and systems thinking, yes. More than ever before.
Honestly, the fear-mongering around AI replacing developers seems to really just expose the developers who never learned architectural and systems thinking, and were just translating Jira tickets to code. While I do not wish job loss upon anyone, I’m not very surprised if those types of jobs will disappear.
If AI is an abstraction, then it's a really shitty abstraction, because i have to dig into lower level all the time to understand what it is doing and tell it to what to optimize.
It's an abstraction layer w.r.t. how you write code, not an abstraction layer in the live system. An IDE is a higher abstraction than e.g. vi but it doesn't mean you don't have to go low-level sometimes.
I think its perfectly fine to have AI write and review code but someone needs to keep track of the business logic. Otherwise its difficult to debug what happens when things break or when context misses a few older assumptions. We go into an infinite loop with LLM writing, reviewing, documenting, designing and validating software one feature after the other but right now, we are context limited. Plus if LLM ever gets stuck, it's hard to prompt it away unless somebody understands the output.
askonomm · · focus · HN ↗
I guess time will tell if the consumer will adapt to the lower quality of products, allowing companies to justify the existence of lazy and incompetent developers, or if the consumer will push back, forcing companies to increase the quality of their developers.
Note: I use AI every day and it is entirely possible to create high quality software with it, so long as you are not lazy and incompetent.
whatever1 · · focus · HN ↗
rfgplk · · focus · HN ↗
whatever1 · · focus · HN ↗
It just gets “reviewed” by an LLM, which will find a nitpick while ignoring the huge fire in the core of the design, force the planner to make even more sloppy code to cover for an irrelevant test case. Rinse old tokens and repeat until you hit limits.
bunderbunder · · focus · HN ↗
For example, I recently got brought in to help with quality on a large-scale system that had been ported to a new platform with the help of coding agents. The project was completed and declared operational in record time, but soon after the business discovered that:
1. The promised scalability improvements did not materialize. Instead, it got worse.
2. Observability had been lost. The telemetry was no longer trustworthy.
3. Users stopped trusting it because it was producing incorrect outputs.
What I ended up discovering was that, while it scrupulously kept existing automated tests passing, any behavior that wasn't explicitly covered by a test was free to change any which way. And there were plenty of small things that weren't explicitly covered. Perhaps because the original authors thought they were so obvious and commonsense that they didn't need one, perhaps because mistakes happen. The why doesn't matter. The point is that reality is messy and imperfect, so giving someone a chance to look at things and think, "Huh, that's funny..." is an essential part of defense in depth.
The real worst part was, this whole replatforming was a huge waste of time, anyway. The improvements they were looking for could easily have been accomplished with some controlled incremental changes to the original system. Mostly just removing a few basic and well-known performance antipatterns.
But way back at the outset, the person in charge of the project asked their agent, "What's the best way to X," and the agent gave them a trendslop answer about how Y alternative technology is more scalable and we should just port to that. It was convincing and they were under intense time pressure to just ship some code because leadership is bought into the AI hype and now has the patience of a 4 year old, so they just went with it.
CamperBob2 · · focus · HN ↗
0c3ca83 · · focus · HN ↗
bitwize · · focus · HN ↗
paganel · · focus · HN ↗
dawnerd · · focus · HN ↗
cat-snatcher · · focus · HN ↗
0c3ca83 · · focus · HN ↗
cat-snatcher · · focus · HN ↗
intrikate · · focus · HN ↗
dawnerd · · focus · HN ↗
Could be argued they were going downhill before but it's a much faster decline since 2023-ish
necovek · · focus · HN ↗
0c3ca83 · · focus · HN ↗
Luajit is under 80,000 lines of code.
whateveracct · · focus · HN ↗
necovek · · focus · HN ↗
If I was to employ them to review the code without giving each the same baseline multi-page prompt, they go into endless loop of "improvement" with no end goal in sight.
More and more frequently, I instruct frontier models to stop and go back to the task at hand.
vandopereira · · focus · HN ↗
[dead]
perrygeo · · focus · HN ↗
Ambolia · · focus · HN ↗
teiferer · · focus · HN ↗
Where does all that supposed productivity go?
al_borland · · focus · HN ↗
Kotlopou · · focus · HN ↗
The problem is that as a lay user of Firefox, I don't know how you could even make it better in terms of features (I could see it being faster etc.).
al_borland · · focus · HN ↗
Powdering7082 · · focus · HN ↗
<a href="https://innovationgraph.github.com/global-metrics/git-pushes" rel="nofollow">https://innovationgraph.github.com/global-metrics/git-pushes
Just because you haven't installed new software doesn't mean that new software doesn't exist.
californical · · focus · HN ↗
nemetroid · · focus · HN ↗
necovek · · focus · HN ↗
teiferer · · focus · HN ↗
legulere · · focus · HN ↗
tripledry · · focus · HN ↗
If I put on my optimist hat for a moment, I hope LLM's would finally make this obvious for people, more stuff !== more value. 250k lines of code a week isn't a flex IMO, it doesn't really mean anything out of context.
perrygeo · · focus · HN ↗
So to answer the original question, "Where does all this supposed productivity go?": Sitting in a github repo somewhere, undeployed.
pretendscholar · · focus · HN ↗
mstaoru · · focus · HN ↗
No new browser, no new iOS clone than runs on Android, no new easy to use DaVinci, no new CAD suite, no $5 SolidWorks clone, no redesigned K8s, no 10x performance speedup in Linux kernel.
bitwize · · focus · HN ↗
A couple more step functions in model capability of the type we've seen in the past year, and there will pretty much be no reason for humans to be involved in the development process at all. All humans would need to do is communicate clearly what needs to be made and flag problems as they come up.
xpct · · focus · HN ↗
pydry · · focus · HN ↗
It didnt go wrong
And if it did, it was because you werent using the latest model.
And if you were, it was because you didnt have the appropriate guardrails.
And if you did, it's because you didnt have AGENTS.MD.
And if you did, it's because you didnt prompt it properly.
And if you did, it you're still going to be redundant soon because I'm sure the next model released will fix whatever went wrong.
mywittyname · · focus · HN ↗
I mean, sure, I could have predicted in what ways an LLM would fuck up, but there's just so many ways I can't keep up.
We just had a major production issue because someone's LLM wrote queries against dev databases. Which are very obviously dev databases because they are labelled with dev in the name, and in the table descriptions. AI reviewer didn't catch it, neither did the human reviewer for that matter.
ysajang · · focus · HN ↗
weatherlite · · focus · HN ↗
Kinda what i'm doing already, but for the young startup I'm at that's surprisingly tons of work. I miss the days we wrote code by hand boy those were fun 8.5 hours workdays.
necovek · · focus · HN ↗
Sounds like the easiest thing in the world: I wonder why did we not think of it earlier?
rgoulter · · focus · HN ↗
Brings to mind this classification <a href="https://en.wikipedia.org/wiki/Kurt_von_Hammerstein-Equord#Classification_of_officers" rel="nofollow">https://en.wikipedia.org/wiki/Kurt_von_Hammerstein-Equord#Cl...
"""I distinguish four types. There are clever, hardworking, stupid, and lazy officers. Usually two characteristics are combined. Some are clever and hardworking; their place is the General Staff. The next ones are stupid and lazy; they make up 90 percent of every army and are suited to routine duties. Anyone who is both clever and lazy is qualified for the highest leadership duties, because he possesses the mental clarity and strength of nerve necessary for difficult decisions. One must beware of anyone who is both stupid and hardworking; he must not be entrusted with any responsibility because he will always only cause damage"""
banannaise · · focus · HN ↗
Now instead of 90% stupid and lazy (harmless, useful for grunt work) you have 90% stupid and hardworking (aggressively causing damage).
automatic6131 · · focus · HN ↗
conmod278 · · focus · HN ↗
pphysch · · focus · HN ↗
danlitt · · focus · HN ↗
Based on what? This will not happen!
DanielHB · · focus · HN ↗
You see this already, LLMs are a lot more reliable in statically typed languages with strong memory guarantees (like typescript or rust) than in weaker languages.
IMO the only way LLM code can avoid most of the pitfalls of human code is if we make new programming languages targeted at being used by LLMs exclusively. Think of languages with very strong methods for formal proofing and stuff like that.
The problem is that even if said language was invented, it would still fail catastrophically when integrated with systems not made in said language. We are very lucky that relational databases already provide a somewhat high level of formal proofing in this regard.
Said language would be impossible to parse by humans, kinda like assembly where you can parse what an isolated piece of assembly code is doing, but if you can't comprehend a somewhat large pure-assembly codebase as a whole.
monkpit · · focus · HN ↗
It’s common advice to wire in deterministic checks to your workflow with LLMs - static languages aren’t inherently better for LLMs, it’s that LLMs produce better code when given deterministic feedback, such as compiler results.
DanielHB · · focus · HN ↗
Of course it is not that simple as a large part of how effective LLMs are is due to training data which is hard to get for a new language.
leptons · · focus · HN ↗
datsci_est_2015 · · focus · HN ↗
djmips · · focus · HN ↗
sp1nningaway · · focus · HN ↗
esafak · · focus · HN ↗
Melkazt · · focus · HN ↗
ben_w · · focus · HN ↗
> I hope you can see the stupidity here if you expect to see any deterministic results at all.
Are you expecting humans to be deterministic in the code they produce?
Thanemate · · focus · HN ↗
ben_w · · focus · HN ↗
And?
The p(that kind of error) is pretty small now. At what point does a probability coming out of an LLM look like "knowing", such that spitting out the wrong answer despite that probability looks like a health problem, a typo, or even just boredom? (Thinking of the Lizardman constant here: <a href="https://en.wiktionary.org/wiki/Lizardman%27s_Constant" rel="nofollow">https://en.wiktionary.org/wiki/Lizardman%27s_Constant)
It's a continuum for both them and us, even if the mechanism is wildly different.
> Making mistakes is not the same as non-deterministic.
i.e. when the dismissal is "non-deterministic" when it should be "Making mistakes", is itself a mistake.
p-e-w · · focus · HN ↗
InsideOutSanta · · focus · HN ↗
ofjcihen · · focus · HN ↗
A human who knows 1+1=2 can still say “3” because they misread the question, misspoke, were distracted, or made some other cognitive error. Likewise, an LLM can output “3” because the generation process selected an incorrect continuation. Those are both errors in producing an answer, not evidence that 1+1 somehow has multiple answers.
So yes, human mistakes and LLM sampling are mechanistically different. If your argument is that LLMs and humans can both make mistakes, then major question here is why are we building out huge amounts of infrastructure at unsustainable spending levels to enable LLMs to make the same mistakes as humans.
InsideOutSanta · · focus · HN ↗
It's not, I'm just pointing out that LLMs won't make that mistake.
You could ask an LLM what 1+1 is, and the number of times it says "3" is so small that it makes no sense to worry about it. It will phrase the response differently each time; that's the nondeterminism. But it won't say "3".
> then major question here is why are we building out huge amounts of infrastructure at unsustainable spending levels to enable LLMs to make the same mistakes as humans.
Yes, if we ignore everything else, that seems like a reasonable question. But let's not ignore everything else, like the fact that LLMs are much more productive than humans and likely already make fewer mistakes than the average programmer.
lolakutty · · focus · HN ↗
I think the disturbing fact is that you can take a frontier model with all the intelligence of humanity, and make it say 1 + 1 = 3, but specifically training for it...
A human with that much knowledge will refuse that attempt. There in lies the difference..
monkpit · · focus · HN ↗
lolakutty · · focus · HN ↗
monkpit · · focus · HN ↗
lolakutty · · focus · HN ↗
ben_w · · focus · HN ↗
Hit them with a stick until they answer as you told them to.
lolakutty · · focus · HN ↗
Obviously, for this purpose, human should not have any feelings (because LLMs don't have), so can't feel pain. Or else the comparison can't work.
ben_w · · focus · HN ↗
The weights are updated depending on if the feedback was positive or negative.
It has a functional effect similar to that which pleasure and pain have with us. Not identical, so far as I know there's not been any reports of any machine learning model that is into BDSM, but for the most part functionally similar.
lolakutty · · focus · HN ↗
thaumasiotes · · focus · HN ↗
Well, that's not true.
<a href="https://www.youtube.com/playlist?list=PLO3a3Ax6Yh6bbtKuxfYBPojj_L2ZBpVsI" rel="nofollow">https://www.youtube.com/playlist?list=PLO3a3Ax6Yh6bbtKuxfYBP...
thaumasiotes · · focus · HN ↗
But this is plainly false. This kind of unforced error occurs all the time.
For example, once when I was in high school I traced an error in my math homework to an intermediate calculation of "2 + 2" as being "3". There was no reason.
What we can say about humans is that, if they know that 1 + 1 = 2, (a) they are unlikely to change their mind about this in any kind of lasting or permanent way, and (b) the rate at which they will mistakenly produce other values for 1 + 1 is very low. But it will happen occasionally, and when it does happen, "they just suddenly decided on the wrong value" is an extremely accurate description of what that looks like.
lkjdsklf · · focus · HN ↗
ben_w · · focus · HN ↗
rhdunn · · focus · HN ↗
The behaviour/output of an LLM is not like that. Ask an LLM to create a dashboard to show games by genre and it will generate different results with each run, and each model/model version produces wildly different results.
ex-aws-dude · · focus · HN ↗
As humans we don't have our memory reset multiple times per day
ben_w · · focus · HN ↗
I've seen humans vote for Brexit, re-elect Trump, ask questions clearly already answered in an FAQ, try to pull on a door labelled "push", and insist on giving me homeopathic silicon dioxide pills* that cost £5** for a 10-12 gram packet.
Continual learning is a difference, but not by itself a reason to care about "deterministic results".
Nor, indeed, correct results.
> As humans we don't have our memory reset multiple times per day
Humans need sleep well before they can read a million tokens' worth of written text. We're more like 300k tokens if you're actually reading and not skimming for 16 hours straight.
Again, different (in soooo many ways), but this isn't a relevant difference when the topic is "deterministic results".
* yes, sand: <a href="https://dailymed.nlm.nih.gov/dailymed/fda/fdaDrugXsl.cfm?setid=97903a05-fd97-420c-836f-b8d73277b683" rel="nofollow">https://dailymed.nlm.nih.gov/dailymed/fda/fdaDrugXsl.cfm?set...
** and that was what it cost in the 90s
ex-aws-dude · · focus · HN ↗
There is no RNG involved when I decide to push vs pull the unlabeled door to my building every morning, it becomes deterministic because its baked into memory
You can put stuff in context to deal with this but you can't do that for everything, its not practical and you would blow the context window
ben_w · · focus · HN ↗
That "because" is an unimportant detail.
You, me, everyone, we all fail tasks with some non-zero probability, including tasks we've done many times before. This is literally why typos happen. We don't always catch the typos we make, this is literally why spellcheck exists. We don't always catch logical errors etc., this is literally why professional writers work with professional editors. We don't always spot errors in code we write; this was why compiler errors have to exist at one level, why unit tests and integration tests exist at another, and why despite automated tests we have QA departments.
Statistically speaking, you as a human will have made more comprehension errors reading that paragraph than even LLMs a few years back.
> There is no RNG involved when I decide to push vs pull the unlabeled door to my building every morning, it becomes deterministic because its baked into memory
There's a lot of random inputs into human behaviour. How much they become random outputs is a similar question as for AI.
Your memories are not "baked" at any point; they're made of and by cells doing chemistry. Your entire body is a trillion small bags of chemicals that communicate with each other by leaking out signalling molecules and electrochemical gradients into each other's water supply, and this goes wrong sometimes.
Small (and apparently random) groups of neurons in your (and all animal) brains briefly enter sleep-like states even while you remain apparently awake. This starts well before you even feel tired.
ex-aws-dude · · focus · HN ↗
For an LLM I have no mechanism to even update the weights
hanifbbz · · focus · HN ↗
rfgplk · · focus · HN ↗
This is all that's needed to actually use LLMs nowadays. How is it a "multiplier" rather than an "equalizer"?
swiftcoder · · focus · HN ↗
Because without the responsible human engineer in the loop, it'll all gradually decay in a cascade of edge-cases. This happens with human written code as well (every "we'll replace this prototype before we ship" you've ever worked on), but with LLMs it happens at 10-100x the rate.
bigfishrunning · · focus · HN ↗
These so rarely get replaced
bcrosby95 · · focus · HN ↗
dnikolovv · · focus · HN ↗
askonomm · · focus · HN ↗
sortoflog · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
CuriouslyC · · focus · HN ↗
In optimizing my game I noticed framerate hitching even after efficient algorithms were in place for expensive stuff, which was caused by shaders not being precompiled consistently or assets not being preloaded in time. The Agent who'd been profiling and optimizing had moved many preloads to a loading screen, which caused a long loading lock, and what it didn't move ahead was loaded and compiled at use, creating slow frames since work was being done on the main game loop.
I instructed the agent to create a speculative pre-warming/compilation priority queue with a per-frame budget, with priority being determined by likelihood signals that the asset or shader will be used soon. Then I had the AI run fully headed games and hunt down causes for frames going over 16ms, and work through them until a batch of games had fewer than 1/1000 frames >16ms and no frames over 60ms after a short initial settling period.
The approach, the metrics, the validation system and the loop were "prompt engineering" above and beyond what I would expect from someone who was merely "vibe coding a game."
zxor · · focus · HN ↗
Zardoz84 · · focus · HN ↗
You will be surprised how many times, catches errores made by the AI coding agent. However,as you point, isn't deterministic. And you can guarantee the end results is 100% fine code
huijzer · · focus · HN ↗
What I in general try to teach the other people about AI: It can be a great tool, but check the results! Especially in the case of engineering: Check and then double check.
ls-a · · focus · HN ↗
[dead]
empath75 · · focus · HN ↗
mjr00 · · focus · HN ↗
Yeah. To me it seems very much like the "use dynamic typing for everything" fad. You had a bunch of junior and/or incompetent developers who went around insisting that type declarations are bad, static typing slows down development, you just code so much faster if everything is dynamically typed. And in the context of a new project, they were totally right. It took a few years for the debt to finally catch up, and people realized that these massive, untyped monoliths they had were unmaintainable. Now the two biggest dynamic languages (Python/JavaScript) are effectively typed languages, because nobody uses their untyped variants for serious work.
Dynamic typing still has great uses -- interactive data exploration, putting together quick scripts (though less relevant with AI...), or even just simple prototypes -- but what we tried to do with it at the start, as an industry, was clearly dumb as hell. I suspect we'll look back in 5-10 years and realize that with some of the stuff we're doing with AI, too. It's already happened with things like Gastown.
paganel · · focus · HN ↗
bdangubic · · focus · HN ↗
necovek · · focus · HN ↗
chanux · · focus · HN ↗
I like to put this as "LLMS give lazy and incompetent developers more runway."
patorjk · · focus · HN ↗
bwfan123 · · focus · HN ↗
ilaksh · · focus · HN ↗
bwfan123 · · focus · HN ↗
How does the system behave in a variety of scenarios including failures and restarts. How is state maintained coherently. There are the kinds of systems problems that an engineer needs to reason through, and if there are bugs in such decisions, they end up becoming costly. I dont expect AI or LLMs to solve these problems at all, since each of them has nuances and tradeoffs which are specific to each system. In short, there is specification complexity in precisely describing system wide behaviors, and unfortunately, there is no lean/tla+ to meaningfully describe systems at scale. You could then ask: How can a system have guaranteed behaviors if they cannot be even stated or proved formally ? The answer to this is how protocols like raft/paxos initially convinced us of their behaviors which is in human review and understanding. That begs the question: How can human review and understanding be reliable, and the answer is that it is not reliable, but humans have ability and processes to continuously learn from experience in the real world. So, our understanding is grounded not only by whats out there in books etc, but also by our own interactions with the world.
Long story short: The responsibility for system-wide behaviors of software systems relies on human review and understanding, which while imperfect can continuously learn.
nottorp · · focus · HN ↗
icedchai · · focus · HN ↗
taurath · · focus · HN ↗
icedchai · · focus · HN ↗
Kuyawa · · focus · HN ↗
Just don't. Fire them! AI is better than a thousand devs. What you need is testers that know what to test that AI can't, not code or UX/UI (not talking about playwright here) but business intelligence if that is testable, the things that produce results (profits) and the reason it was asked for in the first place, to solve a problem
If the problem was asked wrongly, the result will be wrong too. Fire devs, then PMs, then IT Managers if they really don't know how to outperform AI, and that's exactly the point, they won't be able to do it in code or tests or reviews, only in intelligence, for now...
topherPedersen · · focus · HN ↗
One reason I'm reluctant to hand over all of my work to the ai is I don't want to forget how to program or let my skills deteriorate. Another reason is I don't want to become dependent on ai and find myself in a situation where I'm not able to fly/navigate/land the airplane if my auto-pilot or ai malfunctions or fails.
Then the last reason I don't want to take the lazy approach: When I've done "one shot tests" a lot of times the ai will try and take some lazy half-ass shortcut that we would not accept if it were a human doing the work. A lot of times it just doesn't do what you ask it to do.
Where I've found ai extremely helpful though is asking questions about our codebase, or asking it to build me a function that takes in a, b, c arguments and spits out x, y, z.
AI really is one of the greatest things mankind has ever produced, but I don't think it's so good yet that it can replace humans completely. Using it as a form of leverage though I think is what people should be doing. I suppose we'll see what happens to developers who let the ai take over completely. Some people are arguing that if you don't let the ai takeover completely your career is doomed, but personally I think you might be doomed if you forget how to fly the airplane by hand.
askonomm · · focus · HN ↗
We solve business problems through technology and to me it’s quite concerning how many developers think their job was knowing syntax of a particular language. Nobody besides themselves care about the syntax of a particular language, certainly the business doesn’t care. I question if retaining your knowledge of programming languages is that important anymore, but your ability to read code if needed (after all, most programming languages are quite similar so it’s not that hard), do architectural and systems thinking, yes. More than ever before.
Honestly, the fear-mongering around AI replacing developers seems to really just expose the developers who never learned architectural and systems thinking, and were just translating Jira tickets to code. While I do not wish job loss upon anyone, I’m not very surprised if those types of jobs will disappear.
mrheosuper · · focus · HN ↗
cweld510 · · focus · HN ↗
mrheosuper · · focus · HN ↗
another_twist · · focus · HN ↗