‹ BackHN Continuity

Thread

Coding is not solved

584 points · 544 comments · firstSpeaker

  1. temp00345 · · focus · HN ↗
    I read such articles more or less every day. This article would be 100% correct if it came out 1 year ago, 75% correct 9 months ago, 50% correct 3 months ago and it's probably 25% correct now if not less.

    I totally understand where this is coming from. I too am struggling with accepting that my 30+ years of programming experience is quickly becoming obsolete. I'm losing sleep about this, it's tough.

    But just go ahead and give the latest models (Opus 5.5 / Astra 6 as of today) another try. See what they are capable of and read the code which they produce. Any problem area, low level C++ or high level Typescript or Clojure or a weird combination of these..

    Don't be shy, give them a big task, let them build an entire app, UI and all..

    Now compare the output to Opus 4 or gpt-5 from 1 year ago - when they couldn't put together a single function without it being weird and buggy.

    This is exactly my problem, not that the models are very good already, but how fast they got so good. So if coding is not solved yet, it'll get there very soon.

    1. bluegatty · · focus · HN ↗
      You're still doing 75% of what you did before.

      Now the syntax is handled for you, you have a research assistant, and someone that can really dig through the details for you.

      The rest ... is still there.

      1. [deleted] · · focus · HN ↗

        [deleted]

      2. enraged_camel · · focus · HN ↗
        >> You're still doing 75% of what you did before.

        No, I really am not. I'm doing maybe 10% of what I did before. The rest is filled up by other, usually higher order tasks like planning, product management and work orchestration.

        1. bluegatty · · focus · HN ↗
          I'm using multiple agents to do work, with sub-agent orchestration, and it's still mostly coding.

          I feel as though we are all doing the work of 2 people + an Architect, not so much 'Product' issues, although I'm sure it varies.

          But I can't imagine how any actual software is written with 10% of the effort.

          1. datsci_est_2015 · · focus · HN ↗
            Yeah, if anything I feel as if I’m working twice as hard for the same pay. I’m working through 10 “concepts” per week rather than 5 “concepts”. It’s like taking an accelerated university course (Calc 1&2 or Chem 1&2) except it’s every day for the rest of my career.
            1. bluegatty · · focus · HN ↗
              This is the underrated comment of the day.

              Software was always intellectual work, this is Engineering Short Order Cook kind of stuff.

        2. stackbutterflow · · focus · HN ↗
          That's like including your commute as a part of your job. That's not what people are talking about when they say coding is solved.
    2. keel-control · · focus · HN ↗
      Coding is literally solved with current models. It's just a matter of inference cost at this point. Imagine if Astra 6 Max was $0.00001 per 1M ouput...

      I can't think of a single example apart from perhaps the 99.9th percentile difficulty of work that wouldn't be solvable with that configuration.

      1. OliveronData · · focus · HN ↗
        Yet when I wanted the model to implement the naive surface nets algorithm, Astra Max wrote 5+ allocations on inner loops. Alright, fine perhaps, given that I had not told it to preallocate during the inner loops. Except I did give it explicit instructions not to use malloc, and use the arena/pooling API I provided for everything. That was in AGENTS.md.

        Alright, fine, I pointed that out. So what did Astra Max do? It replaced the mallocs with raylib MemAlloc functions, which are malloc wrappers. Why? I had explicitly asked not to use any raylib functions or includes on the module with the prompt. And my AGENTS.md has a minimal raylib inclusion note. Asking isn't helpful because AI does not know anyway. But seriously, why? I asked anyway and Astra said it messed up.

        Fine, I was able to get it on the 3rd try with arenas. But it had added getters for the internal state (guess who had a clause not to make getters?), and when pointed out, Astra Max included the physics header instead and added some convenience functions there, because apparently that was good practice and DRY.

        If I let the reins slip just a little bit, everything turns into a mush; I wish I could get spaghetti instead!

        So whenever I read something like this, "coding is literally solved with current models," I always treat it as a self report. I get that certain section of web frameworks may have been completely RL'd to hell and back, and that React and its ecosystem was already made with the explicit purpose of commodotizing the programmers so any panic 1000 junior hires could write something and maybe even contribute before they get laid off. I get that. But there is a whole world out where reality just doesn't work out like that.

        FYI, this project was in C. C, as in one of the languages that the LLM's should have the highest amount of data for. So if the model had "literally solved" the programming of today, why is it so god awful?

        That was rhetorical. What is solved is the most of the javascript ecosystem being RL'd. Any field that can't be RL'd to that degree (which apparently C isn't), coding is very much not solved.

      2. slopinthebag · · focus · HN ↗
        more like you work on 20th percentile problems and then try to extrapolate to everything
    3. mglvsky · · focus · HN ↗
      but where are new shiny apps from whom who uses $RECENT sota models?
      1. Palomides · · focus · HN ↗
        on the front page of HN multiple times every day? I think it sucks but it's here now
    4. bob1029 · · focus · HN ↗
      > Don't be shy, give them a big task, let them build an entire app, UI and all..

      I have found that a willingness to look like a temporary dumbass (primarily to yourself) is the largest predictor of success with pretty much everything.

      What are the consequences of asking an LLM for the moon and receiving low earth orbit instead? Who cares if the proverbial rocket explodes on the pad? This is all happening entirely in a computer system completely under your control and likely at relatively low cost. No one else has to find out about your mistakes if you don't want them to.

    5. Tade0 · · focus · HN ↗
      > But just go ahead and give the latest models (Opus 5.5 / Astra 6 as of today) another try. See what they are capable of and read the code which they produce.

      Well, for one, they're capable of draining our (or companies') wallets.

      I resolved a huge merge conflict for $60 today. Opus 5.5 did a great job and spent just 1h 16min on this. I could probably run six such sessions today, if I disregarded the need to read and understand the code.

      This money has to come from somewhere and my concern is that it will be from decreasing the number of people hired and/or their salaries.

      At the same time I firmly believe people who had a tendency to produce tech debt will keep doing that, regardless how brilliant LLMs will become. Unscrewing this is going to cost a lot of money.

      1. hajile · · focus · HN ↗
        These companies are operating at huge losses, but are still cutting prices. I don't think companies are ready for what happens when the bills come due and they have to charge enough to not only be profitable, but pay back the years they've been blowing hundreds of billions of dollars.

        Solving a merge conflict for $60/hr is not so bad, but what if it's $600/hr?

        1. Tade0 · · focus · HN ↗
          Interestingly, the merge conflict was caused by a much greater pile of code shoveled in there, meaning the spend is self-sustaining.

          I guess we'll find out after Anthropic's IPO.

    6. lnkl · · focus · HN ↗
      >This article would be 100% correct if it came out 1 year ago, 75% correct 9 months ago, 50% correct 3 months ago and it's probably 25% correct now if not less.

      Feels like I read comment similar to this one each year since 2023.

      1. throwawayffffas · · focus · HN ↗
        I can tell you it was not correct in January of 2026. The real swift started happening with the latest models opus 4.7, fable 5, kimi k3, glm 5.2.

        That's when the models started to be coherent enough for real work.

        They still fuck up, but it does not feel the code was written by drunk interns anymore.

        1. CoolestBeans · · focus · HN ↗
          A month ago I was told January of this year was the inflection point. Month before that the inflection point was December of last year. I'm not saying the tech isn't getting better but are the fundamental limitations being surpassed or are the long tail failures just being pushed further away? Because if its the latter this game of "well models really got good nine months ago" won't stop.
          1. throwawayffffas · · focus · HN ↗
            It looks to me that this generation of models has reached a level of competence where work can be assigned to them and completed satisfactory for various degrees of satisfactory.

            This reflects my own experience with these models. It's not a matter of inflection point if you ask me, it's a matter of accruing capabilities last year the output was not up to my standards 98% of the time, now it looks more like 30% of the time.

            I am sure next year models will be better, but the point where the models begun being good enough to start using seriously for my use cases has now passed.

            1. ModernMech · · focus · HN ↗
              Honestly I feel like 90% of the improvement has been connecting them to deterministic tooling that can tell them when they're wrong. Whenever I chat with one of these latest frontier bots, they still exhibit all the same problems they used to, But throw them in a tight loop with tools that connect to reality, they adjust their behavior accordingly. It's like SLAM vs. dead reckoning for anyone who knows robotics.
        2. its_k1r4 · · focus · HN ↗

          [dead]

    7. roncesvalles · · focus · HN ↗
      I've really tried this and it's not quite what you're saying. You still need to steer it. You still need to stop it from doing stupid things. You still need to vet the architecture and data model. You still need to be clever, it's not that clever.

      I think what software development is turning into is the art of setting up the program into a shape from which the LLM can reliably and efficiently fill up the rest of it. Sometimes that might not even require writing a single line of code. But I think you still need to be as good of a software engineer as you were before.

    8. o_nate · · focus · HN ↗
      We also read these kinds of rebuttals almost every day. The original article made a number of substantive critiques about where AI falls short in the actual requirements for building maintainable, reliable, business-critical software. So its not enough just to point to newer models without explaining how the newer models solve these problems. Can the new models take full end-to-end ownership of a system? If not, how do they solve the problem of humans taking ownership of AI generated code?
      1. [deleted] · · focus · HN ↗

        [deleted]

      2. NCFZ · · focus · HN ↗
        But most software isn’t critical and doesn’t need the level of reliability that’s sacrificed when using LLMs.

        I don’t see the ops comment as a rebuttal. He agrees coding is not completely solved. However, it’s getting closer to being solved.

        1. Panzer04 · · focus · HN ↗
          I don't think I really agree.

          If anything, the core job of most software is to reliably and repeatedly do the thing a user wants. Bugs outside eof that are OK, but a software that doesn't Do The Thing is useless.

          LLMs increasing bug density in exchange for more features is likely a terrible tradeoff for most software people use IMO.

      3. aunty_helen · · focus · HN ↗
        Yes they can and tomorrow the models will be better. By the time code hits prod and is used in anger the frontier will have moved.

        It’s really not an acceptable opinion to have, that one day we’re going to be developers drowning in ai generated code. That’s a problem for opus 6 in 4-5 weeks time.

        Developers are still very much needed, business people just don’t have the chops to wire a deterministic system together. But acting like some poor class naming or 2k loc file is the disaster it once was is dated.

        1. fleischhauf · · focus · HN ↗
          can't be sure the development continues. As much as opus got better at programming the less I understand the output. it keeps inventing terms and skipping steps for the reader to understand what is written properly. I also feel like progress in programming is becoming a lot slower now.
    9. Mabusto · · focus · HN ↗
      This is where I’m at as well. A friend recently showed me what opus 5.5 is capable of and it was both awe inspiring and disturbing.

      I’ve been using AI as a great pair programming partner for about a year now, but every time I prompt it to write code agentically, it just makes a pile of crap that I end up spending more time fixing. This is so very quickly becoming not the case anymore.

      I think core engineering skills are never going away (I’ve been doing this for 25 years and come from a background doing C++ for video games), that you will always need to have a mental model of the code and if you’re going to call yourself “professional” you need to be able to go in and fix/build by hand. But you’re going to get left behind if you aren’t at least willing to eat some humble pie and re-evaluate your views on agentic coding every few months as these tools get better.

      Grieve, I know I have, but there is joy on the other side. I used to love getting lost in flow state with Soma.fm playing and the phone unplugged, that is still there, but what that looks like is changing fast and table stakes in this industry has been “adapt or die” for as long as I’ve been in it.

      1. rhines · · focus · HN ↗
        I'm using Opus 5.5 at work extensively, have also used Kimi 3 and Deepseek 4.1 for stuff outside of work, and I can't deny that the tools aren't capable. They solve problems, surface issues I wouldn't have considered, and largely do write better code than I do.

        The one exception is for stuff that I care very deeply about. For a very small subset of projects where I'm willing to spend hundreds of hours ensuring that what I build is the best possible thing, AI still hasn't been able to match my work unless I micromanage it, but at that point just writing the code myself is actually faster.

        For almost all jobs though AI is probably the future, most software developers never cared that much about their company's code anyway.

    10. oulipo · · focus · HN ↗
      Well probably you haven't done so yourself, because I use them on a fairly high-complexity system with concurrency, IoT devices, etc, and the solutions are often in need of large rework almost every time (when touching core stuff)
    11. throwawayffffas · · focus · HN ↗
      I concur as well in the last six months we moved significantly in the it just works direction. Eight months ago, all output was either garbage or weird and buggy at best. Today the models actually solve problems and fix bugs.
    12. an0malous · · focus · HN ↗
      I use the latest models every day for app development, it does still make many mistakes and poor decisions. Just today I dealt with an issue where it allowed a silent failure that would delete all of the users data without anyone noticing. The code it writes for one task is usually pretty good, but it still doesn’t always follow conventions well even if you’ve documented them. It also has no sense for long-term architecture or organization, or when to make tradeoffs for less complexity because it’s a temporary or prototype feature that needs to stay easy to change. I’ve found that if I do even 3 months of pure agentic coding with no code reviews, it’s aged 10x faster than a human coded codebase so it’s like a 30 month legacy codebase now. I know people who are obsessed with AI coding and they’re throwing away projects they’ve built over a year and starting over because it’s become too slow to make changes.

      What kind of coding are you using these models for? Most of the people I know who share your perspective never go beyond the prototyping stage. I’d be curious to hear from anyone who’s been AI coding for more than six months, shipping it to real users, and isn’t looking at their code at all.

      1. ysajang · · focus · HN ↗
        I'd be curious to hear from anyone who's been AI coding for more than six months, shipping it to real users, and isn't looking at their code at all.
        1. spicyusername · · focus · HN ↗
          What do you want you hear?

          I haven't opened my IDE in months. I only skim the MRs to make sure there are no glaring issues.

          Success is a good set of skills, a good up front design for your feature, architecture docs in your repo, lots of automated linting, unit testing, AI MR reviews, and black box e2e testing.

          I doubt I'll ever _have_ to write a line of code again.

          All of my time is spent juggling 5-20 agents, mostly planning and responding to the agents escalations.

          GPT 6 Sol / GPT 6 Luna are my bread and butter and spec-driven design works great.

          Like 75% of the small issues that used to take me a few hours or half a day in 2021 take me like 15 minutes to ship now start to finish, and I can ship them while doing other things.

          1. datsci_est_2015 · · focus · HN ↗
            Context matters here. Are you a freelance website designer? Are you a CRUD app codemonkey? Do you work for a glorified data integrator, maybe in the healthcare industry? Are you an embedded engineer? Do you write code that if something goes wrong someone will die?

            I can also ship slop if the stakes are 0. The code I write has stakes. Thousands of dollars per false positive, which happens to be approximately the entire margin that we’re operating within. If a false positive occurs and I can’t explain it because I shipped slop, then it’s curtains for me - and our contract.

            1. maxnevermind · · focus · HN ↗
              I think "March of 9s" helps to describe this situation. Some devs seems to be happy with just one 9 as it suffice for their use-case while for others one 9 means that they would loose client's data 10% of the time or some other catastrophic consequence.
      2. gantam · · focus · HN ↗

        [dead]

    13. lbrito · · focus · HN ↗
      I've been reading such arguments more or less every day for years now: dude, you need to try the latest model. Forget about last week's model, it didn't work. This week's model is the real deal.
    14. whateveracct · · focus · HN ↗

      [dead]

    15. [deleted] · · focus · HN ↗

      [deleted]

    16. mrheosuper · · focus · HN ↗
      By that logic, the software by the top AI-token maxxer(Microsoft, meta, Twitter) will be much, much better than ever.

      Is Windows much better than they were 2 years ago, why didn't MS just point Astra/Fable(or whatever SOTA, non-public model they have access to) to Windows kernel and ask it "make it 10x faster" ? Are they stupid ?

      1. furyofantares · · focus · HN ↗
        This is like when two years ago people said "oh yeah if these things are accelerating software development then where's all the new software??!!"

        The answer then was "it still takes time, man".

        That's the answer now, too!

        It took time for them to accelerate the quantity of software, which is the thing these are most naturally good at, and it will take even more time for them to improve the quality of software, which has proven to be less of a superpower for LLMs than quantity is.

        1. [deleted] · · focus · HN ↗

          [deleted]

        2. mrheosuper · · focus · HN ↗
          Are you saying Microsoft does not have CI/CD pipeline ? If no, where is the "delivery" ?
          1. furyofantares · · focus · HN ↗
            Sorry, but it's unclear to me how that question is relevant even related to the conversation.
            1. mrheosuper · · focus · HN ↗
              MS obviously setting up CI/CD on their product, that means whatever the AI was outputting, should have reached user hand by now. Yet we see no improvement in quality
    17. maxnevermind · · focus · HN ↗
      > But just go ahead and give the latest models (Opus 5.5 / Astra 6 as of today) another try.

      I find my self going opposite direction recently. I start using frontier models less and less and smaller models more and more. I recently tried Opus Opus 5 with ultra-code level of effort and I wasted like 2 days on trying to migrate some project to new libs because it spits out a huge wall of text as replies/plans and waste huge amount of time on things I didn't ask it to check and see no point in checking. At the end I just took a smaller model and used it in more controlled fashion playing the role of orchestrator/planner. Smaller models are good enough and also much faster. My collisions from that, frontier models and high reasoning effort levels are for pure vibe-coding when you don't really control or understand what is happening. My application area is not pure SWE though, I didn't churn a lot of code before and I don't now, so I have enough time on understanding/polishing my projects if needed.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.