‹ BackHN Continuity

Thread

The Claude Delusion

97 points · 164 comments · hn_acker

  1. muglug · · focus · HN ↗
    > Of course, the more you know about a subject, the less convincing the AI's responses are.

    This is said all the time by AI skeptics and I think it's right in some areas and massively wrong in others.

    I know (or at least assume I know) a lot about certain coding domains where frontier models also show convincing ability. And we know that frontier LLMs really do excel in some areas of mathematics (i.e. when an inexpert human was able to prompt the models to derive a closer bound on the Riemann Hypothesis).

    OTOH I know those same models struggle to do things I'm not an expert in (e.g. writing English in a captivating way) because I read their output and have taste.

    1. howunfortunate · · focus · HN ↗
      I think part of this comes from the fact that LLMs are surprisingly good at logic but roughly about as good as expected on information accuracy.

      LLMs are not convincing to me in the domain I did grad school...but neither is Wikipedia, or Reddit, or random pop sci books. And LLMs are basically just summarizing those things.

      But when made to work through difficult arbitrary logic (like coding), they are very impressive.

      I think this also explains why people gripe a lot about LLM coding _style_, but concede that LLMs do totally fine on coding _correctness_ in 2026

      1. QuantumGood · · focus · HN ↗
        More perhaps at the appearance of logic. I still find all models make easy to find mistakes in logic, if you think carefully about what they say. Usually they allow "close enough" assumptions that are not, actually, close enough.

        When I ask about acoustics, they still often make incorrect assumptions, e.g. overlooking that they are speaking of logarithmic display of digital levels when the discussion has shifted to SPL (sound pressure level in air).

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.