‹ BackHN Continuity

Thread

The Normalization of Inexplicable Failures

277 points · 124 comments · pxx

  1. pmarreck · · focus · HN ↗
    I am big on reproducibility (nix aficionado) and determinism (flagging test failures are a red-alert, all-hands-on-deck situation in my world) and correctness.

    I am also big on testing (the correct things). And nine-nines (big on Elixir).

    And... I'm also big on agent-assisted dev. Which requires pretty much every check in the book to stay productive in. And that's fine to me. I've seen bugs that I wouldn't have made myself. And I've also seen my own bugs fixed. They've all gotten fixed in short order. I don't see why this is a problem.

    Raise your personal standards.

    Thing is, the unreliable-software situation was already untenable before agents (in poor hands) made it worse.

    1. lokar · · focus · HN ↗
      I don’t think the author (or many people) doubt that one can (and some will) find a way that does not “suck”

      But it’s pretty clear that most people are not. For whatever reasons (mgmt pressure, trying to get ahead, skill issues, etc) they half ass it, accept the 10% (silent) fail rate and blame the bad outcomes on the AI as if that absolves them. Or, adopt the attitude that 10% fail is fine, and people who say otherwise are being picky, or are anti-ai luddites or whatever. You should accept that things will suck.

      1. gchamonlive · · focus · HN ↗
        If you took a bad but functional AI generated service and transported it back to 2018 it would have been at worst just mediocre. People do seem forget how dreadful devslop was in the past. I'd take an AI generated mess to disentangle every time over a spaghetti codebase that grew organically in the hands of careless managers.
        1. majormajor · · focus · HN ↗
          I don't like the idea of shaming people for making OSS vibe-coded tools for niche hobbies, so I don't want to name names, but there is ABSOLUTELY some stuff on Github now that can be used to get the job done but has UI/API/code performance, consistency, and quality issues, that would've been unfathomable for the average OSS project in 2018. Because it's the sort of stuff that only happens when there isn't a human in the loop to point out some very-obvious swings-and-misses. Like "you don't need a third button here doing the same thing as these other two" or "this button literally does nothing in 3/4 of the modes, but it's never shown as disabled" or "this takes 5 times longer than it needs to and blocks the main UI thread because work is all happening sequentially."

          In the past it wasn't really common at all to add 10 features in an evening without actually trying to use those 10 features by hand yourself.

          Personally I think it's wonderful that tools for these spaces exist when they didn't use to. But it's also ludicrous to say that anyone with Claude can replace even mediocre homemmade stuff in any dimension other than "being worth building even a bad tool" or "getting to semi-usable faster." Currently you still benefit massively from knowing what's going on behind the scenes, and from knowing "software engineering 201" type stuff around what sort of testing would be helpful where, vs accepting model-default-output everywhere.

          1. gchamonlive · · focus · HN ↗
            > there is ABSOLUTELY some stuff on Github now that can be used to get the job done but has UI/API/code performance, consistency, and quality issues, that would've been unfathomable for the average OSS project in 2018

            I will take time to address you comment fully, there are lots to unpack, but I'd like to stop here and point out that comparing "absolutely some stuff" from today to the "average 2018 project" might be unfair. If you are going to compare the bottom of the pile of synthetic code, you should do the same with organic code of back then lest you draw an unfair comparison.

            1. johnnyanmac · · focus · HN ↗
              > If you are going to compare the bottom of the pile of synthetic code, you should do the same with organic code of back then lest you draw an unfair comparison.

              Make it the bottom 70% AI vs. the bottom 10% human if you want to. You're still going to be well within a sea of AI-slop because the volume is that large. The big thing about bad human code is that it tends to still be compact; I can read in some minutes the intention and what does(n't) work.

              For each vibe-coded project I gotta do a tiny expedition just to get the basic idea of what's going on. let alone figuring out if things actually work as intended.

              1. gchamonlive · · focus · HN ↗
                > The big thing about bad human code is that it tends to still be compact

                I envy the environment you've worked before, because you definitely had a better experience than mine. My first job was working on a product -- not a prototype you see, an actual product with paying customers generating over 100k USD monthly for the company -- that was completely designed by interns from the ground up. Once we received a pull request on a part of the system that dealt with calculation reports that made snapshots of the relational db into MongoDB. This application was in PHP using an outdated CakePHP framework -- which was outdated for a good reason since it used to introduce breaking changes in minor releases (citation needed, but who's got the time these days...) -- and the merge request that once-employee left us to figure out how to merge was a single 400 lines function. You can extrapolate from it to infer the quality of the snapshots we were taking and you probably wouldn't land too far from the horrors we've seen.

                So pardon me, but my experience shocks violently with your affirmation.I think you underestimate what the bottom 10% really is.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.