‹ BackHN Continuity

Thread

The LLMentalist Effect (2023)

235 points · 309 comments · jalev

  1. atemerev · · focus · HN ↗
    All of this is accurate (and useful). But yeah, psychics do not prove theorems from frontier math and build working complex software.
    1. Madmallard · · focus · HN ↗
      Neither does LLMs. Not without a monumental amount of human oversight and expertise.

      Everyone not in a particular field asking LLM about said field is rolling bad dice.

      1. atemerev · · focus · HN ↗
        I have published a LLM-assisted proof of a long-standing problem in convex geometry (70 years) while not being a specialist in convex geometry. I am a scientist and I have education and scientific experience in computer science and systems biology, but I was never a mathematician. So, well, you can definitely work across fields at least.
        1. tripledry · · focus · HN ↗
          Do you have a link? I would also be very interested in seeing the prompts, from what I've seen you still need to understand the field, maybe I'm wrong.
        2. Madmallard · · focus · HN ↗
          Okay? Maybe that's a specific use-case where they can possibly make progress. It's also likely that most of the leg-work for that specific proof has already been done and just a bit of recombination of leading theories and publications resulted in the answer.

          If you have it ask for involved legal documents to give to a lawyer like 100% of the time they find problems with it. And someone who isn't in law would have not known any better.

          When you have it write complex code that is not easy/quick to test, especially things that are specifically NOT concretely defined, like net-code (because it's all on the trade-offs you want to accept for your particular game), it's going to just repeatedly create sync issues.

          I've asked it to write-up a detailed explanation of the different types of turns in 4-panel dance games and it's just permanently wrong no matter what I say.

          The more unique and lacking of training data that exactly represents the problem statement the more impossible the statistical machine will generate text that makes sense.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.