‹ BackHN Continuity

Thread

The Claude Delusion

97 points · 164 comments · hn_acker

  1. muglug · · focus · HN ↗
    > Of course, the more you know about a subject, the less convincing the AI's responses are.

    This is said all the time by AI skeptics and I think it's right in some areas and massively wrong in others.

    I know (or at least assume I know) a lot about certain coding domains where frontier models also show convincing ability. And we know that frontier LLMs really do excel in some areas of mathematics (i.e. when an inexpert human was able to prompt the models to derive a closer bound on the Riemann Hypothesis).

    OTOH I know those same models struggle to do things I'm not an expert in (e.g. writing English in a captivating way) because I read their output and have taste.

    1. keeda · · focus · HN ↗
      I can tell whether AI responses in subject matter I am (or was) not familiar with are correct when I can "make them work" for me, i.e. whether they actually solve my problem, and on the whole they have been pretty solid.

      Coding of course is something I know well and it obviously works very well for me there, but computer vision was not, and it has helped me solve lots of useful, bespoke problems because I can very literally "see" if it works.

      However, it has even helped me solve problems in arbitrary matters far outside my expertise. Choice example: a complicated multi-airline, multi-jurisdiction flight delay compensation case which companies like AirHelp turned away, and ChatGPT got me literally hundreds of dollars that neither airline was willing to hand out. It told me what to say to whom and why, and when I said it, the responsible airline capitulated.

      What could be more convincing than cold, hard $$$?

      The trick of course, is to figure out how to validate the output of the AI, which can take some effort on our part. But many would rather downplay the technology than take an honest crack at making it work for them, and I suspect it's because they're already prejudiced against it or their incentives are otherwise not aligned.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.