‹ BackHN Continuity

Thread

What About Rails?

334 points · 236 comments · jrochkind1

  1. bionsystem · · focus · HN ↗
    As an SRE I would be very interested on experienced devs point of view on that stance, "we don’t even necessarily need to read the code the LLMs produce". To me, that is the only way a single dev can manage > 1 agent. Because I feel running the code will always be slower than a single agent generating it. On the other hand, it implies lack of human understanding on what is going on under the hood. Which is fine if you trust the LLM to write great code, and great tests for the code, but fundamentally you have to have 100% trust, 99.9% is not going to be enough in any serious industry, would it ?

    Also eventually you'll also have to trust it to write the deployment code or even run the deployment itself, otherwise SRE is going to be the bottleneck. And only then should I feel anxiety about the rest of my career (that, or my employer decide LLM are good enough to get rid of me, even if they are imperfect).

    1. viraptor · · focus · HN ↗
      > Which is fine if you trust the LLM to write great code, and great tests for the code, but fundamentally you have to have 100% trust, 99.9% is not going to be enough in any serious industry, would it ?

      This is lacking a lot of nuance. There are many types of code. There are many situations where I'm analysing something one-off and if I get 33% success ratio, but can easily verify the result, I'm happy - still saved me time and money. They're are situations where I'm generating graphs from some dataset and I don't have to trust anything - I know what the result should look like, I just need the agent to drive matplotlib. There are low stakes dashboards which I'm happy to generate and develop entirely via agents - they'll embed the updated screenshots in PRs that I can yolo-merge - worst case is that someone complains about something not working next time they visit. Then there's lots of experimenting which was never stopped to hit production anyway.

      Finally after all of that you get code that's actually part of deployable features. Of course the trust is nowhere near 100%, but if you have a healthy testing process (e2e, validating different browsers, or whatever is appropriate for your environment), then what's your trust in human developer+review? Because mine is nowhere near 100% either.

      In practice there are places where I extremely don't care about the code and never wanted it anyway, places where I'll read the code to check the design or just in case, and places which agents are not allowed to touch (medical billing rules for example).

      1. weaksauce · · focus · HN ↗
        dhh did just what you did and apparently made a bunch of mistakes in his powerpoints.
        1. viraptor · · focus · HN ↗
          I don't understand what you mean by "just what you did". I'm not making PowerPoint presentations.
          1. weaksauce · · focus · HN ↗
            vary similar processes. vibe code/vibe create a presentation. he also knew what "shape the graph should look like" but still had errors in his powerpoint. if you can't trust it with a powerpoint i'm not sure why you can trust it for other low stakes things.
            1. viraptor · · focus · HN ↗
              Have you got any sources? The only thing I can find is the 2026/2006 mix-up which feels like a classic typo rather than AI mistake. There are other typos there which makes it less likely the presentation was fully generated. And it means that the result wasn't reviewed - looking at code is irrelevant in comparison.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.