‹ BackHN Continuity

Thread

The Four Horsemen of Agentic Coding

114 points · 89 comments · haute_cuisine

  1. JV00 · · focus · HN ↗
    Disagree on the slop point. By now the code of frontier models works well, and the need to read it is gone. The Claude word salad problem was pretty much solved with the 5.5 opus release.

    The other 3 points made in article make sense though.

    1. thethirdone · · focus · HN ↗
      If you don't need to read the code at all, what are the engineers doing to make the other issues matter? Deskilling and alienation are still about the relationship between engineers and the code.
      1. ramoz · · focus · HN ↗
        You still have to orchestrate these things and the wholistic system architectures, pull the slot machine handle many times to get the right thing, and review/verify the working output constantly.

        That said, I think we need to be prepared for how a simple prompt like "make it faster" will work in a relatively short timeframe - and where the AI is able to adequately refactor/re-platform/be done in a way where even the architecture becomes obscure to us.

        1. SpicyLemonZest · · focus · HN ↗
          I often use simple prompts like “make it faster” and get back functionally correct commits. When I review them, I often discover that they didn’t do what I really wanted - sometimes Claude made it faster by incurring expensive costs, or doing a migration that would be risky in prod, or making a tradeoff that’s not worth it.

          Perhaps I could set up a workflow to handle this without reading the code. I’ve seen people who make Claude fill out arcane templates, and then spend their time reviewing reports and tinkering with loops or whatever. But why would I want to read a TPS report about the code rather than just reading the code?

    2. v64 · · focus · HN ↗
      > By now the code of frontier models works well, and the need to read it is gone.

      Like anything else, it depends on what you're working on. For small personal projects I agree, but in maintaining large scale business applications with millions of lines of code, I am still having to guide the models to reuse existing code and to be more flexible with their data structures to accommodate changing business requirements in the future.

      They only see a snapshot of the world the application exists in due to their ephemeral nature, and until that is overcome, they will only be able to learn general principles and not the particular nuances of your codebase that only arise from observing how the users use it on a daily basis.

    3. formerly_proven · · focus · HN ↗
      Really seems more like you got used to their verbal tics and unmonitored quality, to be honest.
      1. JV00 · · focus · HN ↗
        DHH, antirez, steipete all said something along the same line, about not needing to read all code anymore. Are they used to unmonitored quality too?
    4. fwlr · · focus · HN ↗
      The Claude word salad was certainly dialed back somewhat, but there are still a few flies in the soup, so it’s far from solved.
      1. JV00 · · focus · HN ↗
        Ok, if when it happens you just tell it, it does rewrite it better. Not always human coworkers can do that :)
        1. fwlr · · focus · HN ↗
          Yes, there’s always flies in the soup, and sometimes they just serve you a whole plate of insects, but you can’t beat the price. Just ask the waiter to get you a soup with less flies, and keep doing that until the number of flies is acceptable to you.

          How about, no thanks, I’ll go to a real restaurant.

    5. tschellenbach · · focus · HN ↗
      I don't know why people are downvoting this. It depends on what you're building.

      AI build me a cost overview of xyz: works perfectly without ever looking at the code

      AI build me unreal engine 5 level graphics, or a database, or infra at high scale: this will need some active steering for now :)

      1. JV00 · · focus · HN ↗
        And even for active steering, does it require reading the code? Iteration can be done on the result of the work itself, abstracting code away
      2. smashed · · focus · HN ↗
        They will do these prompt just fine but the result will be a near identical copy of an open source version. And it won't tell you it cloned some existing engine, because it does not know, it just pulled it out from its model. Now you have an AI slopped 3d engine or DBMS no-one cares about that will just rot.
    6. wonnage · · focus · HN ↗
      “Solved” = roughly back on par with 4.8. This is like Apple rolling back the butterfly keyboard and telling you it’s the best keyboard yet
      1. JV00 · · focus · HN ↗
        No come on. 4.8 was borderline unusable. I get intelligible results from 5.5
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.