‹ BackHN Continuity

Thread

One Month Without AI

188 points · 228 comments · saibotk

  1. stack_framer · · focus · HN ↗
    My employer pays for Claude, and my approach is to use it as a better Google search. It's often not better.

    Just today it made three glaring mistakes in one session:

    1. It read a file in the wrong directory, because that file had the same name as the file in the right directory. It apologized when I challenged it, promising me that it would remember to "read import statements" in the future.

    2. It miscounted the number of times a function was called in my repo. It said 20, while my built-in IDE search accurately showed 17. Again, it apologized when I corrected it.

    3. It referred to a variable by name that does not exist anywhere in my code. It apologized, and said it was referring to a variable used internally by one of the third-party packages installed in my repo.

    So many apologies.

    It's the little things like this that remind me on a regular basis just how little I can trust artificial "intelligence."

    1. pydry · · focus · HN ↗
      I'm amused at how many hacker news accounts saw this and immediately jumped to the conclusion that you either were using the model wrong or using the wrong model.

      ive been regularly seeing this exact reaction online since November 2025, sadly :/

      1. dumberquestions · · focus · HN ↗
        It's just a little baffling to see someone describe a level of performance I haven't experienced since 2025, despite frequently using the tech, as being a frequent concern.
        1. binary0010 · · focus · HN ↗
          Might be the type of projects you are working on and how much you care about performance and code quality.

          Working on more complex, logic heavy projects with strong performance needs I find the models to be useful but certainly not 'one shot' on pretty much anything. And often incredibly frustrating and genuinely bad code that collapses performance and bloats systems - like what a really bad junior might write.

          When I'm working on large standard crud web projects with already decent architecture and a good harness and skills, honestly they work pretty well a lot of the times, the code looks good, does what it needs and fits in with the architectural style.

          I really think a lot of hn people just write simple repetitive software, and a small portion works on complex, weird, dense, and novel'ish logic projects. The two obviously don't have the same experiences.

          1. [deleted] · · focus · HN ↗

            [deleted]

        2. Mawr · · focus · HN ↗
          To a first approximation, since 2025 was at best a year ago, it's not reasonable to be so surprised that the improvements in that time weren't enough for some use case. There's only so much anything can improve in a year.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.