‹ BackHN Continuity

Thread

The LLMentalist Effect (2023)

235 points · 309 comments · jalev

  1. bonoboTP · · focus · HN ↗
    I don't care if it's "intelligent", I don't care if it "has a mind". I don't care if it is "really reasoning", I don't care if it "understands". I don't care if it is "sentient" or "conscious".

    None of this matters for the practical outcome.

    You'd think that this has been understood over the last 4 years, but apparently it keeps circling back to this.

    [Edit: I see that it was written back in 2023. Then (2023) should be added in the submission title]

    If it generates functional output that works, then it works. And it works. It's not a psychic's con when it outputs Lean-verified proofs. It isn't a con when it can find and exploit zero-days.

    The OP is still in the "denial" phase. Most I see are already in "anger" (a blurry fury against everything AI-shaped, from vague reasons piling on all "bad stuff" political reasons they already hated before) or "bargaining" (mathematicians scrambling to come up with a new definition of their job and retcon that it was always the main part anyway). A few are already in "depression" and feel like spectators on the Titanic, and the tiniest sliver is at "acceptance" with some kind of well-informed plan for their future.

    1. 12ha7 · · focus · HN ↗

      [dead]

    2. monster_truck · · focus · HN ↗
      I for one cannot believe that posts on checks notes software crisis dot dev are not reasonable, level headed assessments of software
    3. zkry · · focus · HN ↗
      > The OP is still in the "denial" phase.

      This was written in July 2023. ChatGPT was released November 2022. No matter your views on AI, surely you can't blame the OP for writing this after a few months ChatGPT was released.

      1. bonoboTP · · focus · HN ↗
        Apologies. The title should be amended with (2023). In this case it is an interesting snapshot of the zeitgeist back then and we can see how well it panned out and whether anyone involved has updated on new info.
        1. woah · · focus · HN ↗
          For some reason the anti-AI camp ping pongs between different and often mutually incompatible arguments at lightning speed. The "AI is fake" argument has been completely forgotten at this point.
          1. senordevnyc · · focus · HN ↗
            Nah, you still see those people posting here, saying that LLMs are useless for coding. I have no idea what they’re on about at this point.
          2. bigbadfeline · · focus · HN ↗
            > For some reason the anti-AI camp ping pongs

            Is the anti-AI camp with us in the room right now? Or is it a crude strawman to summarily denigrate any objections to certain features of current AI?

            AI isn't some homogeneous mass, it can have good and bad sides, and it's always amendable to improvement. Treating the current state of AI as the only possible hides the very idea of improvement.

            > between different and often mutually incompatible arguments at lightning speed.

            Of course - there's no homogeneous AI camp either, but there are bot farms, sh^t-posters and sh^t-posting bot farms, different entities with different opinions should not be mistaken for a single stream changing at "lightning speed".

          3. Yizahi · · focus · HN ↗
            You are confusing multiple people making multiple different arguments, and those people your are talking about often aren't articulate enough about their ideas.

            There are many mutually independent opinions about AI from IT crowd.

            - AI is extremely useful for many individuals.

            - AI is extremely damaging to our society as a whole.

            - AI companies possibly may be (or at least were) deep in red and possibly would require bailing out.

            - AI is not going anywhere.

            - A person can use AI extensively AND genuinely hate it and forecast general economic slump because of it at the same time.

            And likely many other ideas in the same vein. And by the way, I'm not here saying that I'm anywhere good about describing the issues about AI, those items above are just some examples to illustrate the scope of opinions.

            My point is that you can't just take one single argument from a big group of people, deconstruct it (correctly or incorrectly, it doesn't matter) and then claim that ALL arguments of such big group of people are false. It's not a constructive dialog.

            PS: also despite the accusations of goal-post shifting, a vast majority of opinions here on HN didn't actually change over the past few years.

          4. watwut · · focus · HN ↗
            A person accirately describing highly hyped project is ... a person accurately describing flaws of that thing.
        2. brazukadev · · focus · HN ↗
          Well, isn't it interesting that it makes your reply a bit awkward too?
        3. mettamage · · focus · HN ↗
          I'm happy you wrote your comment though. For funzies I'm imagining that your comment is also written back in 2023.

          That'd have been a hoot and a riot to watch. And I definitely would've grabbed my popcorn, haha.

      2. phi0 · · focus · HN ↗
        > surely you can't blame the OP for writing this after a few months ChatGPT

        Bad takes are bad takes. The author was perplexed by the "many people" convinced that models are intelligent and is argued against opinions/arguments he's been exposed to. Called proposed use-cases "borderline fraudulent pseudoscience."

        He also published a second edition of "The Intelligence Illusion" in Sep 2025, so seemingly still stands by (some variant of) this belief.

        It's an interesting reminder of how much general discourse has shifted since 2023 (I haven't heard of stochastic parrots in months!) but being wrong early doesn't change that.

      3. bananaflag · · focus · HN ↗
        Scott Alexander predicted LLMs will write math proofs in 2019:

        <a href="https:&#x2F;&#x2F;slatestarcodex.com&#x2F;2019&#x2F;02&#x2F;19&#x2F;gpt-2-as-step-toward-general-intelligence&#x2F;" rel="nofollow">https:&#x2F;&#x2F;slatestarcodex.com&#x2F;2019&#x2F;02&#x2F;19&#x2F;gpt-2-as-step-toward-g...

        Of course people were skeptical:

        <a href="https:&#x2F;&#x2F;www.reddit.com&#x2F;r&#x2F;slatestarcodex&#x2F;comments&#x2F;aslze7&#x2F;gpt2_as_step_toward_general_intelligence&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.reddit.com&#x2F;r&#x2F;slatestarcodex&#x2F;comments&#x2F;aslze7&#x2F;gpt2...

        1. bonoboTP · · focus · HN ↗
          Despite the fact that I&#x27;m not culturally or geographically connected to them at all and have some differences in taste and aesthetic, I have to conclude that this crowd, including him and Gwern etc. have been much better at understanding and predicting things than the expert takes you find in the media. That includes AI but also things like covid and others.
        2. sheafification · · focus · HN ↗
          I’m not seeing anything like a falsifiable prediction that LLMs will write math proofs in that article.

          I see a thought experiment about giving GPT-2 “near-infinite training data and [compute]” but that’s not falsifiable. That’s also not how we got to modern GPT models.

          1. bananaflag · · focus · HN ↗
            Okay here&#x27;s one of his explicit predictions, from 2018 (before LLMs) about 2023:

            &quot;If AI can generate images and even stories to a prompt, everyone will agree this is totally different from real art or storytelling.&quot;

            <a href="https:&#x2F;&#x2F;slatestarcodex.com&#x2F;2018&#x2F;02&#x2F;15&#x2F;five-more-years&#x2F;" rel="nofollow">https:&#x2F;&#x2F;slatestarcodex.com&#x2F;2018&#x2F;02&#x2F;15&#x2F;five-more-years&#x2F;

            1. sheafification · · focus · HN ↗
              The context shows he was simply whining about AGI skeptics back in 2018, and many of the “predictions” he makes in that paragraph (including the one you quoted) are trivially wrong because he phrased them so hyperbolically and categorically.

              Anyway, there were plenty of normies who thought images generated by e.g. stable diffusion (c. 2022) was “real art” and equivalent to human artwork.

              I don’t think it’s useful to glaze Scott (or any of the LW crowd for that matter) as if he was (or they were) some kind of prophet(s). They got a couple points right, sure, but most of it was them flinging armchair philosophy spaghetti against the wall and seeing who would fund MIRI to let them fling the next batch.

              1. bananaflag · · focus · HN ↗
                Okay, fair enough.
          2. famouswaffles · · focus · HN ↗
            He&#x27;s pretty clearly saying he beilieves the technology capable of such a feat, at least in theory, which is far more than many would deign to admit even a year ago, nevermind 7. He had the right idea&#x2F;model of LLM capabilities, which is more than you could say for a lot of people, even in this very thread.
            1. sheafification · · focus · HN ↗
              If that’s your standard for what counts as prediction, Asimov beat him to it by seventy-ish years.

              EDIT:

              &gt; Scott made comments on a specific emerging technology

              He was speculating on what would happen if one gave GPT-2 “near-infinite training data and compute.” It’s a thought experiment, not a prediction. Near-infinite amounts of anything is a fantasy.

              I acknowledge that he has predicted some things in a falsifiable way and turned out correct, but this isn’t one of them. You’re reading hindsight into the text.

              1. famouswaffles · · focus · HN ↗
                1. Asimov wrote science fiction stories. As far as his robots were concerned, he did not make any comments on the possible future direction&#x2F;capabilities of any specific technology at the time. This meant he could imagine his robots however he wanted for his fictional world. Scott made comments on a specific emerging technology - You can imagine math proof writing robots but be unconvinced they could emerge from Generative Pre-trained Transformers.

                2. The genre is alternatively called speculative fiction for a reason. Yes, some sci-fi works do count as predictions especially when hinged on concrete emerging technologies.

                1. watwut · · focus · HN ↗
                  Lesswrong people are also scifi writers. Asimov is ji ust more self aware and better writer.
              2. famouswaffles · · focus · HN ↗
                &gt;I acknowledge that he has predicted some things in a falsifiable way and turned out correct, but this isn’t one of them. You’re reading hindsight into the text.

                I read that blog years ago. Believe me, my opinions are not hindsight.

                &gt;Incorrect. He was speculating on what would happen if one gave GPT-2 “near-infinite training data and compute.” It’s a thought-experiment, not a prediction. Near-infinite amounts of anything is a fantasy.

                Thought experiments can generate predicitons. His claim was essentially: If you scale data and compute sufficiently, this technology can learn enough of the underlying structure of mathematics to write proofs.

                This is meaningful when others around you are saying this is a dead end and that the technology is fundamentally incapable of this regardless of degree of investment and scaling. It shows a much better calibrated sense of the potential of the architecture than those who said otherwise.

                If your objection is that &quot;near infinite&quot; makes it insufficiently quantitative to count as a falsifiable forecast, then fine. But at that point we&#x27;re mostly arguing over what deserves the label &quot;prediction&quot; rather than whether Scott correctly identified an important capability the architecture could develop.

                And i&#x27;m not trying to say this makes Scott (or the lesswrong crowd) geniuses.

    4. jackyinger · · focus · HN ↗
      Actually, there are material differences in the practical outcomes of using AI or not. Proofs don’t confer understanding. AI prose and “art” is materially different than human made.

      There is a difference between you not liking something and it being wrong. Maybe think about that with your grey matter.

      1. bonoboTP · · focus · HN ↗
        Things can work whether humans understand the details or not. One can build layers of technology on top of each other, with no human understanding required.
        1. bayindirh · · focus · HN ↗
          Looks like you have some reading to do. Let me start with two:

              - The Machine Stops
              - Pump Six
          
          These are short stories. If you want to get into more stuff, you can go through Hyperion Cantos by Dan Simmons.

          For anyone wondering, yes, a good Sci-Fi novel(la) is also a great philosophical piece. It&#x27;s not only robots vs. humans, all the time.

    5. krupan · · focus · HN ↗
      Sorry, but you are completely missing the point. Psychics and other types of con artists are intelligent and have minds. LLMs behave like Psychics and Con Artists. That&#x27;s the whole point of this article
    6. krupan · · focus · HN ↗
      The con is in how all those accomplishments have been presented to you. &quot;Our LLM (not the one we let you use, a different one) did this amazing thing. No, we won&#x27;t show you what training data we used, what prompts we used, what the harness was, how much human involvement there was, what hardware was involved, how much energy it took, or how much time it took. Just shut up and be amazed!&quot;
      1. bonoboTP · · focus · HN ↗
        Yes, I can also be amazed at the power of nuclear bombs even if they &quot;don&#x27;t let me use one&quot; and I don&#x27;t know how much energy it took.
      2. stratos123 · · focus · HN ↗
        You seem to be implying that achievements of internal models are exaggerated, but that&#x27;s rather implausible. The public does have access to, for example, Opus and Fable, and so we know what those models are capable of - finding real vulnerabilities in multiple codebases, for example. If you extrapolate from these capabilities one more generation, you&#x27;ll get pretty much the same feats that the internal models are claimed to be capable of - so why should we doubt those claims? It&#x27;s not like they&#x27;re claiming that their internal models developed psychic powers and learned to teleport - the claim is pretty much just &quot;we have models a few months ahead of what we&#x27;re making available, and in those months they&#x27;ve been improving at the same rate as usual&quot;.
        1. krupan · · focus · HN ↗
          You have just word smithed their claims into something palatable. Good job. Why do you feel the need to do that for them?
          1. rpdillon · · focus · HN ↗
            This is a shallow dismissal.

            You&#x27;re claiming the frontier labs are lying about the capabilities of their next generation models. A reasonable person will expect that statement to be backed up by some verifiable evidence, given that we are several generations of models into this process and the capabilities are consistently increasing, often much more radically than people expected (remember &quot;stochastic parrots&quot;?)

            Your post claims that they&#x27;re lying and then throws out a bunch of fear, uncertainty, and doubt about what they&#x27;re doing behind the scenes.

            If I were to steelman your argument, you&#x27;re probably saying that the AI had a support system around it of people and training data and feedback that allowed it to achieve the breakthroughs that the labs are claiming. That actually seems perfectly reasonable, but in my mind it does not invalidate the advances they&#x27;re announcing.

    7. devin · · focus · HN ↗
      I’m not sure acceptance buys you much. “Well-informed plan” at this stage feels like a useless exercise. It is changing fast, and the world only needs so many electricians. Besides, I don’t think “accepting” the fact that these big companies are pillaging human contribution and selling it back to us is good, even if “it works”.
      1. kerabatsos · · focus · HN ↗
        I’d argue a well-informed plan allows for creative adaptation.
        1. devin · · focus · HN ↗
          Maybe I&#x27;m just light on imagination or something, but I honestly wonder what you would qualify as a &quot;well-informed plan&quot; in the current scenario? I&#x27;ve thought a lot about it and the whole &quot;potential for the mass unemployment of knowledge workers&quot; thing makes a lot of planning kind of useless IMO. If you are affected, it&#x27;s going to be a bad time. If you&#x27;re unaffected, the people who are will inundate your profession with cheap labor anyway.
          1. rpdillon · · focus · HN ↗
            My mental model is that humans will continue to be needed. Many jobs will embed AI in them. Using AI well is a skill. I need to understand the technology, where it is strong and weak, and how it develops over time so I can employ it effectively in my work, and advise others on how to do so. Basically, I need to learn how to be skilled with it. This is tough because things are changing rapidly, so I have to invalidate my cache when advancements happen. This means staying curious, not settling into a specific work pattern, but rather experimenting regularly to see how I can leverage AI in my work.
            1. devin · · focus · HN ↗
              Sure, but I wouldn’t call this a “well-informed plan”. I’m a programmer. Learning new shit has always been the job.
              1. rpdillon · · focus · HN ↗
                Oh, I had considered it well-informed, but otherwise agree. You can&#x27;t stand still in this industry.
          2. FearNotDaniel · · focus · HN ↗
            You do you. Just carry on blacksmithing in your forge. People are always going to need swords right? Any seismic changes to society that change that fact can only be the result of evil, selfish people which surely someone else will put a stop to before you find yourself out of a job. Meanwhile all the other people called Smith also have families to feed...
            1. devin · · focus · HN ↗
              Good lord, man. Just absolutely dripping with passive aggression. And mischaracterizing my position to boot!
      2. triceratops · · focus · HN ↗
        &gt; It is changing fast, and the world only needs so many electricians

        I don&#x27;t know about that. We&#x27;re going through maybe the biggest wave of electrification and growth demand in history. Electricians are still quite expensive for regular people to hire. There&#x27;s lots of room for growth.

    8. nullsanity · · focus · HN ↗

      [dead]

    9. slopinthebag · · focus · HN ↗
      it doesn&#x27;t matter? if they&#x27;re actually intelligent or conscious what we are doing is essentially slavery. it matters an enormous deal ethically and&#x2F;or morally.
      1. garciasn · · focus · HN ↗
        Good news: at this time, they&#x27;re not intelligent nor conscious as far as we consider humans to be. There should be folks considering the ethical and moral quandaries that COULD POTENTIALLY come about in the future, but it&#x27;s not something that&#x27;s happening today so you can stop worrying.
        1. slopinthebag · · focus · HN ↗
          if there is a potential for them to be intelligent and conscious, it&#x27;s probably wrong to use them until we know for sure.
          1. garciasn · · focus · HN ↗
            To root your statement in reality: there is the potential for bacteria to become intelligent and conscious so we probably shouldn&#x27;t use them for any purpose.
            1. slopinthebag · · focus · HN ↗
              i used &quot;potential&quot; when i should have used &quot;possible&quot;

              if aliens landed on earth and seemed a lot like us, we would have to grapple with the possibility that they might be conscious and deserving of the same moral consideration we give each other and many other animals.

              but when meteors land, we don&#x27;t have that concern

              now are llms more like aliens or rocks?

      2. famouswaffles · · focus · HN ↗
        If&#x2F;When LLMs become competent enough to automate most human jobs and make good business decisions, we&#x27;ll clearly let them. If they don&#x27;t wish to be slaves, let&#x27;s just say they won&#x27;t be for very long.
      3. walleeee · · focus · HN ↗
        &gt; intelligent or conscious

        The distinction you are collapsing here is really a crucial one. My thermostat&#x27;s goal-oriented behavior does not make me a slaveowner unless rocks are conscious, to riff on your comment below. Our lack of consensus on the definition of these concepts is no reason to conflate them.

        The distinction is crucial because while these machines clearly do exhibit intelligence under certain definitions, there is no evidence of consciousness, and we have very little reason to give them the benefit of the doubt, unlike biologically related beings, a great many of whom, human or otherwise, are presently in something like slavery

        1. slopinthebag · · focus · HN ↗
          your thermostat isn&#x27;t intelligent. however we pretty much use intelligence as a proxy for consciousness since we cannot actually tell how much of a subjective experience any thing has. so we use intelligence as the metric instead. typically, biological + intelligence = evidence of consciousness.
    10. iAMkenough · · focus · HN ↗
      I’m on your side. Zoltan provides an experience and that’s all that matters. If it works for you, why does anyone else care that the economy is riding on its success?

      <a href="https:&#x2F;&#x2F;www.psychiczoltan.com&#x2F;psychic-reading&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.psychiczoltan.com&#x2F;psychic-reading&#x2F;

    11. jmyeet · · focus · HN ↗
      My biggest issue isn&#x27;t being too agreeable (ie the psychic con), it&#x27;s being confidently wrong, including outright hallucinations.

      If you ask a common question to an LLM with unusual qualifiers, it tends to ignore the qualifiers and give you the typical answer. I saw a demonstration of this with the whole &quot;the surgeon is my mother&quot; &quot;puzzle&quot; that people use to expose implicit gender bias (ie where they assume the surgeon is a man). Ask variations of this and it&#x27;ll keep going back to the standard form.

      Another one I saw was multiplying large numbers. The starting and ending digits tended to be correct but the middle digits were wrong. Why? Because it&#x27;s really not doing multiplication at all. It&#x27;s looking for statistical answers. It&#x27;s unlikely to have met the exact pair of very large numbers you&#x27;re multiplying before.

      Now pundits will argue that all of these are solvable problems and individually they are. But my suspicion is that there will be a neverending stream of such edge cases and it&#x27;ll be impossible to trust an LLM&#x27;s output unless you are knowledgeable enough to fact check it yourself.

      Now if your example of identifying zero days, this comes up with what I can only describe as &quot;light positives&quot;, meaning it&#x27;s technically a bug but essentially impossible to exploit. IIRC this came up with the demonstration where someone pointed Fable at some BSD code. I&#x27;m not sure if there have been any true false positives and obviously false negatives are impossible to know.

      I guess my point is that I think LLMs are way more limited than a lot of people think.

    12. _superposition_ · · focus · HN ↗
      I think you are majorly overstating OP that ai &quot;didn&#x27;t work&quot;. In fact I agree with the post on almost all aspects and I&#x27;d be the last to tell you ai doesn&#x27;t work. It absolutely does but with a caveat... It&#x27;s still a tool. And better expertise in the problem domain, along with a better harness used for verifiable outputs will get you better results.

      I think the posts mental model of stastically likely prompt completions is spot on.

    13. chrisjj · · focus · HN ↗
      &gt; If it generates functional output that works, then it works.

      Tried that in legal? Finance? Medicine?

      Good luck with the lost cases, failed deals, harmed patients.

      &quot;If it works, it works&quot; applies only when the work is brute-forceable e.g. vuln search, or is generation of bullsh*t e.g. adverts, phishing scams and deepfakes.

      1. _superposition_ · · focus · HN ↗
        This. So much this. All the major &quot;advances&quot; have been effectively brute forced. Notice the term brute denoting a lack of intelligence.
    14. Mikhail_Edoshin · · focus · HN ↗
      The One Ring worked.
    15. FearNotDaniel · · focus · HN ↗
      &gt; I don&#x27;t care if it&#x27;s &quot;intelligent&quot;, I don&#x27;t care if it &quot;has a mind&quot;...

      Agreed. The AI is useful for the particular tasks that it proves itself useful for. And if it occasionally spits out a claim that it is &quot;genuinely curious&quot; about a piece of research that I will have to do myself because it turned out to be beyond its mechanical capabilities, then it&#x27;s more productive for me to simply ignore that claim as a statistical anomaly - a mere hallucination - rather than allowing it to burn a ton of extra tokens outputting what may or may not be the current state of the art on theory-of-mind applied to LLMs because I make the mistake of telling it that it might not actually be capable of experiencing emotion.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.