From personal experience, frontier models absolutely struggle with understanding a physical situation based on words. (Okay, I haven't played with Astra much. GPT-5.6 Sol makes outrageous errors that anyone understanding a real world object would not make. And I was just asking it about NPT threads, not advanced physics.)
But seriously, what's up with these benchmarks? The example question in the paper is:
> PHYBench, problem 140: equivalent expressions for the same rope tension
> Problem statement. Three identical homogeneous balls are placed on a smooth horizontal surface, touching each other and are close enough to each other. A rope is wrapped around the spheres at the height of their centers, tying them together. A fourth identical sphere is placed on top of the three spheres. Find the tension T in the rope. It is given that the weight of each sphere is P.
For some reason the paper was focused on the fact that the grader didn't notice that some models were producing answers that were trivially algebraically equivalent to the reference answer. But this is missing the elephants in the room:
1. "touching each other and are close enough to each other": the right response is "hey, Professor, what do you mean 'close enough to each other'? They're sitting on a table in an equilateral triangle, all touching (i.e. tangent at their equators), right? Did you have a different configuration in mind?"
2. The answer is 0. Go find four baseballs or foursquare balls or whatever, make a little triangle with three of them, and balance the fourth one on top. It's not especially hard on an appropriate surface. Now loosely wrap an imaginary rope around them (but see below) to keep them from moving - no tension is needed because they're not moving anyway. So the models and the reference answer are wrong, IMO.
3. How, exactly, do you plan to wrap a rope around the spheres, at equator height, with no built-in tension (not pre-stretched), such that the rope does not immediately fall off? Friction? But I suspect you need to pretend there is no friction to get the reference answer. (Or maybe that the marble-marble interface has friction but the marble-table interface doesn't? Again, I haven't tried to reverse engineer it.) So maybe the right answer is "infinity or impossible -- in the scenario where the rope is needed, the rope will promptly fall off because it cannot be stable in the described configuration and gravity pulls it down, and once the rope falls off the tension will be zero and the top marble will fall and the other three will roll over the rope."
4. The answer might be "any tension you like -- just wrap the rope with the desired amount of tension". Imagine three baseballs in a triangle with a rubber band around them and a fourth baseball on top for good measure. The tension is a function of what rubber band you choose.
I'm sure there's an interpretation of the question that makes the reference answer correct, and I was not inspired to try to reverse engineer it.
My tentative conclusion is that LLMs are almost unbelievably good at solving problems that are fully contained within the inputs and (training/verification) outputs, and that they and the people training them are not actually particularly good at the input and output parts. If you are training a model to benchmaxx this benchmark, you are training a bad model.
I'm also a trained physicist, spent some time in industry, and taught a college freshman math class for a semester a couple decades ago.
There's an interpretation that makes the reference answer correct. I didn't realize this until halfway through teaching the math class: Based on the lectures, textbook, and homework problems, you match the "form" of the problem with similar problems in the textbook, and apply the same algorithms to solve it. The students have some vague idea of this, referring to it as finding the "trick," but it's never explicitly explained to them.
I was glad for the interesting puzzles and good grades. But physics really came alive for me in the lab, where mother nature decides the conditions of the problem being solved.
> Based on the lectures, textbook, and homework problems, you match the "form" of the problem with similar problems in the textbook, and apply the same algorithms to solve it.
Is the ultimate goal here to train an LLM to do well at mediocre physics homework or to train it to be a high-quality tool that can do real physics?
And yes, I'm well aware that, even in elementary school, this form-matching is a thing. It's kind of sad.
The odd thing is that we do train people to do real physics, somehow despite my cynical take. I see the education process as following heuristics that are believed to have some relationship to developing real world abilities, even if we don't know why. We all worked these problems, and now we're somehow able to do physics. In the end we don't know what turns people into physicists. Or musicians, artists, etc.
An analogy is making students write 5-paragraph essays. We don't really believe that the knowledge of how to write a 5-paragraph essay is a real-world ability, and the LLM's can write them all day long. But for some reason we believed that 5-paragraph essays were a good pedagogical tool.
But whether the same heuristics apply to teaching an LLM is anybody's guess. If they learn differently than we do, then feeding them on our learning tools isn't necessarily going to help them.
amluto · · focus · HN ↗
From personal experience, frontier models absolutely struggle with understanding a physical situation based on words. (Okay, I haven't played with Astra much. GPT-5.6 Sol makes outrageous errors that anyone understanding a real world object would not make. And I was just asking it about NPT threads, not advanced physics.)
But seriously, what's up with these benchmarks? The example question in the paper is:
> PHYBench, problem 140: equivalent expressions for the same rope tension
> Problem statement. Three identical homogeneous balls are placed on a smooth horizontal surface, touching each other and are close enough to each other. A rope is wrapped around the spheres at the height of their centers, tying them together. A fourth identical sphere is placed on top of the three spheres. Find the tension T in the rope. It is given that the weight of each sphere is P.
For some reason the paper was focused on the fact that the grader didn't notice that some models were producing answers that were trivially algebraically equivalent to the reference answer. But this is missing the elephants in the room:
1. "touching each other and are close enough to each other": the right response is "hey, Professor, what do you mean 'close enough to each other'? They're sitting on a table in an equilateral triangle, all touching (i.e. tangent at their equators), right? Did you have a different configuration in mind?"
2. The answer is 0. Go find four baseballs or foursquare balls or whatever, make a little triangle with three of them, and balance the fourth one on top. It's not especially hard on an appropriate surface. Now loosely wrap an imaginary rope around them (but see below) to keep them from moving - no tension is needed because they're not moving anyway. So the models and the reference answer are wrong, IMO.
3. How, exactly, do you plan to wrap a rope around the spheres, at equator height, with no built-in tension (not pre-stretched), such that the rope does not immediately fall off? Friction? But I suspect you need to pretend there is no friction to get the reference answer. (Or maybe that the marble-marble interface has friction but the marble-table interface doesn't? Again, I haven't tried to reverse engineer it.) So maybe the right answer is "infinity or impossible -- in the scenario where the rope is needed, the rope will promptly fall off because it cannot be stable in the described configuration and gravity pulls it down, and once the rope falls off the tension will be zero and the top marble will fall and the other three will roll over the rope."
4. The answer might be "any tension you like -- just wrap the rope with the desired amount of tension". Imagine three baseballs in a triangle with a rubber band around them and a fourth baseball on top for good measure. The tension is a function of what rubber band you choose.
I'm sure there's an interpretation of the question that makes the reference answer correct, and I was not inspired to try to reverse engineer it.
My tentative conclusion is that LLMs are almost unbelievably good at solving problems that are fully contained within the inputs and (training/verification) outputs, and that they and the people training them are not actually particularly good at the input and output parts. If you are training a model to benchmaxx this benchmark, you are training a bad model.
analog31 · · focus · HN ↗
There's an interpretation that makes the reference answer correct. I didn't realize this until halfway through teaching the math class: Based on the lectures, textbook, and homework problems, you match the "form" of the problem with similar problems in the textbook, and apply the same algorithms to solve it. The students have some vague idea of this, referring to it as finding the "trick," but it's never explicitly explained to them.
I was glad for the interesting puzzles and good grades. But physics really came alive for me in the lab, where mother nature decides the conditions of the problem being solved.
amluto · · focus · HN ↗
Is the ultimate goal here to train an LLM to do well at mediocre physics homework or to train it to be a high-quality tool that can do real physics?
And yes, I'm well aware that, even in elementary school, this form-matching is a thing. It's kind of sad.
analog31 · · focus · HN ↗
An analogy is making students write 5-paragraph essays. We don't really believe that the knowledge of how to write a 5-paragraph essay is a real-world ability, and the LLM's can write them all day long. But for some reason we believed that 5-paragraph essays were a good pedagogical tool.
But whether the same heuristics apply to teaching an LLM is anybody's guess. If they learn differently than we do, then feeding them on our learning tools isn't necessarily going to help them.