‹ BackHN Continuity

Thread

The Normalization of Inexplicable Failures

277 points · 124 comments · pxx

  1. adamddev1 · · focus · HN ↗
    Excellent post. People always defend agentic/LLM-driven development by saying, "Well it's good enough", or "It works most of the time."

    That may be tolerable for some user-facing app. But what if we start normalizing failures in the libraries, the infrastructure, and the compilers? Everything descends into a mess of unreliability, and that slows EVERYTHING and EVERYONE down.

    1. Root_Denied · · focus · HN ↗
      Banking/Finance is the one industry I've seen push back against this type of thinking. Transactions must be handled in a perfect and repeatable way, or the system is unusable as far as the company is concerned.

      There's definitely still AI/LLM integration happening, but is kept out of specific areas of the business.

      1. paulddraper · · focus · HN ↗
        I’ve had banking transactions fail a number of times for unknowable or ill-defined reasons. One just last week in fact.

        So that is a strange choice for repeatable, understandable operations. Might as well use Jev.

        1. Root_Denied · · focus · HN ↗
          From a consumer perspective there's an unknown number of layers between your actions or instructions and what the backbone tech of the bank is doing, but having worked in the space I can point to a few things.

          Firstly, the transaction failed and notified you about the error - that's certainly intentional.

          Second, there's failures that would be invisible to you as the customer, such as "instead moving $100 from account A to account B it credited account B but didn't debit account A, without generating an error". Those are mostly the systems I'm talking about being insulated from AI development. Without more detail about the exact problem you had it's hard to tell if it's a failure in customer facing systems or backend infra.

          Third, I'm assuming you were able to reach out to the bank directly and resolve the issue by talking to a person (if the issue was really outside the norm), which isn't something you can assume will be possible with a lot of customer service ops these days (or you're going to be waiting hours/days for that callback).

          Fourth, you can't really know the error rate of the bank's systems, or how common a given particular error is. It may be a known issue, or it may be a completely unreported one. Assuming it's an error with something on the backend/backbone of the bank's operations, it's running code that can be inspected, reviewed, understood, and fixed - sometimes by a very expensive COBOL consultant.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.