‹ BackHN Continuity

Thread

How good are frontier models at physics?

100 points · 50 comments · qt31415926

  1. RomanKornev · · focus · HN ↗
    Starting to feel more and more like chinese room experiment

    The models are confidently answering physics questions, treating it as a math problem, but they don't fundamentally "get it" and even recently failed simple "should i drive to car wash" test

    The sample efficiency is just crazy low

    Still surprising that even with this they managed to saturate the benchmarks

    1. red75prime · · focus · HN ↗
      > they don't fundamentally "get it"

      There's no clear decision criteria for this. Do trick questions demonstrate that most people don't "get it"? And, well, older model saying dumb things doesn't establish a general principle that LLMs don't "get it" in general.

      > The sample efficiency is just crazy low

      Autoregressive pretraining requires huge amount of data to go from a blank state to a somewhat functional model. Fine-tuning, LORA, reinforcement learning of foundation models and in-context learning are much more sample efficient.

      > Chinese room

      ...creates a wrong intuition that by cranking a Leibniz's mill you are somehow responsible for whether it understands something or not.

      1. rogerrogerr · · focus · HN ↗
        > Do trick questions demonstrate that most people don't "get it"?

        "Should I walk to the car wash" is hardly a trick question. If a human told me to walk to the car wash because it's so close, I would say that demonstrates they don't "get it".

        1. red75prime · · focus · HN ↗
          I don't see that much difference with "A plane crashes on the border of the United States and Canada. Where do they bury the survivors?"

          Fast thinking fails to activate slow thinking.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.