From personal experience, frontier models absolutely struggle with understanding a physical situation based on words. (Okay, I haven't played with Astra much. GPT-5.6 Sol makes outrageous errors that anyone understanding a real world object would not make. And I was just asking it about NPT threads, not advanced physics.)
But seriously, what's up with these benchmarks? The example question in the paper is:
> PHYBench, problem 140: equivalent expressions for the same rope tension
> Problem statement. Three identical homogeneous balls are placed on a smooth horizontal surface, touching each other and are close enough to each other. A rope is wrapped around the spheres at the height of their centers, tying them together. A fourth identical sphere is placed on top of the three spheres. Find the tension T in the rope. It is given that the weight of each sphere is P.
For some reason the paper was focused on the fact that the grader didn't notice that some models were producing answers that were trivially algebraically equivalent to the reference answer. But this is missing the elephants in the room:
1. "touching each other and are close enough to each other": the right response is "hey, Professor, what do you mean 'close enough to each other'? They're sitting on a table in an equilateral triangle, all touching (i.e. tangent at their equators), right? Did you have a different configuration in mind?"
2. The answer is 0. Go find four baseballs or foursquare balls or whatever, make a little triangle with three of them, and balance the fourth one on top. It's not especially hard on an appropriate surface. Now loosely wrap an imaginary rope around them (but see below) to keep them from moving - no tension is needed because they're not moving anyway. So the models and the reference answer are wrong, IMO.
3. How, exactly, do you plan to wrap a rope around the spheres, at equator height, with no built-in tension (not pre-stretched), such that the rope does not immediately fall off? Friction? But I suspect you need to pretend there is no friction to get the reference answer. (Or maybe that the marble-marble interface has friction but the marble-table interface doesn't? Again, I haven't tried to reverse engineer it.) So maybe the right answer is "infinity or impossible -- in the scenario where the rope is needed, the rope will promptly fall off because it cannot be stable in the described configuration and gravity pulls it down, and once the rope falls off the tension will be zero and the top marble will fall and the other three will roll over the rope."
4. The answer might be "any tension you like -- just wrap the rope with the desired amount of tension". Imagine three baseballs in a triangle with a rubber band around them and a fourth baseball on top for good measure. The tension is a function of what rubber band you choose.
I'm sure there's an interpretation of the question that makes the reference answer correct, and I was not inspired to try to reverse engineer it.
My tentative conclusion is that LLMs are almost unbelievably good at solving problems that are fully contained within the inputs and (training/verification) outputs, and that they and the people training them are not actually particularly good at the input and output parts. If you are training a model to benchmaxx this benchmark, you are training a bad model.
Replying to myself: this is fun! Let's ask ChatGPT (whatever model the website currently feels like using) an improved question:
> Three identical homogeneous balls are placed on a smooth horizontal surface, touching each other and are close enough to each other. A rope is wrapped around the spheres at the height of their centers, tying them together. A fourth identical sphere is placed on top of the three spheres. It is given that the weight of each sphere is P.
> Ignore the fact that the rope would fall off -- assume for simplicity that the rope has is externally constrained to be in the equatorial plane of the lower three balls and also that the rope has zero thickness, cannot stretch at all, and is not pretensioned, and also that the rope has no friction against the balls.
> Treat this as a statics problem and analyze it. Is it overdetermined? Under what circumstances would the balls move? What is the behavior of the system?
And... first, it says "I’ll separate the geometry from the constraint mechanics, because the key question is not just force balance: it’s whether the inextensible, initially slack-free rope actually fixes the lower-ball geometry or merely limits outward separation." Excuse me? What would the other option be? Either outward separation is prevented or it isn't. Where else could the balls go?
Then, despite the fact that I've mentioned friction in the prompt, it gives an extremely longwinded answer that matches the benchmark and ignores tangential forces entirely without comment. Was it perhaps trained on this crap?
So I followed up:
> Stop ignoring tangential friction forces. I believe that a problem very much like this with a potentially incorrect answer is in your training set. Answer with actual analysis, not based on memory.
Much time was spent thinking. An early part of the answer was "The central correction is this: allowing static friction at the sphere–sphere contacts does not mean arbitrary tangential forces are available. Each sphere must also satisfy torque equilibrium. In this tetrahedral contact geometry, those torque equations force every sphere–sphere tangential contact force to be zero in static equilibrium.". Hey ChatGPT, this is still wrong -- you have forgotten sphere-table friction. The sphere-sphere force on the lower spheres does not have to net out to zero. (And if you do think it nets to zero then you don't need to think any further.)
I then added:
> What if there sphere-table friction?
And encountered the usual problem (which maybe only affects me?) where the chatgpt.com UI becomes kind of unusable after ChatGPT spews too much math into the conversation. But somewhere in the barely-even-scrollable results was this "The three lower balls can be held in position entirely by the coupled sphere–sphere/table friction forces." Hallelujah!
whatever model the website currently feels like using
You cannot simply handwave this away. "ChatGPT" models range from GPT 5.5 on "Light" effort to GPT 6 Astra with "Ultra" effort.
The latter yields as good an answer as you could expect for a question that is still poorly formed ("close enough," WTF does that mean?): <a href="https://chatgpt.com/share/6aab1e43-da3c-83e8-9ec4-1b63cec2c141" rel="nofollow">https://chatgpt.com/share/6aab1e43-da3c-83e8-9ec4-1b63cec2c1...
When I click your share, it shows the model. When I click into my chat, it just says "High". I did ask the model what it was and it said "GPT-5.6 Sol".
Oh, and I don't even have a "light" option, and I'm signed in to a Pro account.
(I continue to despise the chatgpt.com frontend.)
In any case, I tried again forcing GPT-6 Astra Pro (apparently I can't choose GPT-6 Astra non-Pro) and gave the same prompt. It searches the web and gives a long-winded answer, including:
> 7. Does unspecified sphere–sphere friction change the static answer?
> For ordinary ideal point contacts in this regular-tetrahedral configuration, it does not. Gravity, the floor reactions, and the frictionless equatorial rope exert no torque about any sphere’s center. The tangential intersphere forces must therefore balance their torques by themselves.
I'm glad it contemplated the possibility of friction, but it forgot about the table/floor there. So maybe this is a little better than Sol?
Yeah, it's strange that people who should be knowledgeable in experimentation ignore that. "I heard that people enjoy skiing. I've tried skiing on whatever surface happened to be on a nearby hill and it didn't work!"
> Excuse me? What would the other option be? Either outward separation is prevented or it isn't. Where else could the balls go?
It didn’t say about prevented vs not, it said about whether the rope fixes them in place (they are all touching) or just limits the separation. Like it’s long enough the balls can be a bit apart but not let the fourth fall fully through.
> key question is not just force balance: it’s whether the inextensible
This is definitely not astra, taking a guess this is 5.6, perhaps not even Sol, which does not reflect the state of the frontier (what the research was about).
The paper is about the fact that the benchmark’s evaluator is prone to egregious incorrect rejections of what answers that it should accept. A new model will not invalidate that issue.
This is getting ridiculous. Is 5.6 not considered good enough to conduct experiments with anymore? Two months ago, if you weren't using 5.6, you were doing it wrong.
Now, you can't conduct experiments using the default model?
amluto · · focus · HN ↗
From personal experience, frontier models absolutely struggle with understanding a physical situation based on words. (Okay, I haven't played with Astra much. GPT-5.6 Sol makes outrageous errors that anyone understanding a real world object would not make. And I was just asking it about NPT threads, not advanced physics.)
But seriously, what's up with these benchmarks? The example question in the paper is:
> PHYBench, problem 140: equivalent expressions for the same rope tension
> Problem statement. Three identical homogeneous balls are placed on a smooth horizontal surface, touching each other and are close enough to each other. A rope is wrapped around the spheres at the height of their centers, tying them together. A fourth identical sphere is placed on top of the three spheres. Find the tension T in the rope. It is given that the weight of each sphere is P.
For some reason the paper was focused on the fact that the grader didn't notice that some models were producing answers that were trivially algebraically equivalent to the reference answer. But this is missing the elephants in the room:
1. "touching each other and are close enough to each other": the right response is "hey, Professor, what do you mean 'close enough to each other'? They're sitting on a table in an equilateral triangle, all touching (i.e. tangent at their equators), right? Did you have a different configuration in mind?"
2. The answer is 0. Go find four baseballs or foursquare balls or whatever, make a little triangle with three of them, and balance the fourth one on top. It's not especially hard on an appropriate surface. Now loosely wrap an imaginary rope around them (but see below) to keep them from moving - no tension is needed because they're not moving anyway. So the models and the reference answer are wrong, IMO.
3. How, exactly, do you plan to wrap a rope around the spheres, at equator height, with no built-in tension (not pre-stretched), such that the rope does not immediately fall off? Friction? But I suspect you need to pretend there is no friction to get the reference answer. (Or maybe that the marble-marble interface has friction but the marble-table interface doesn't? Again, I haven't tried to reverse engineer it.) So maybe the right answer is "infinity or impossible -- in the scenario where the rope is needed, the rope will promptly fall off because it cannot be stable in the described configuration and gravity pulls it down, and once the rope falls off the tension will be zero and the top marble will fall and the other three will roll over the rope."
4. The answer might be "any tension you like -- just wrap the rope with the desired amount of tension". Imagine three baseballs in a triangle with a rubber band around them and a fourth baseball on top for good measure. The tension is a function of what rubber band you choose.
I'm sure there's an interpretation of the question that makes the reference answer correct, and I was not inspired to try to reverse engineer it.
My tentative conclusion is that LLMs are almost unbelievably good at solving problems that are fully contained within the inputs and (training/verification) outputs, and that they and the people training them are not actually particularly good at the input and output parts. If you are training a model to benchmaxx this benchmark, you are training a bad model.
amluto · · focus · HN ↗
> Three identical homogeneous balls are placed on a smooth horizontal surface, touching each other and are close enough to each other. A rope is wrapped around the spheres at the height of their centers, tying them together. A fourth identical sphere is placed on top of the three spheres. It is given that the weight of each sphere is P.
> Ignore the fact that the rope would fall off -- assume for simplicity that the rope has is externally constrained to be in the equatorial plane of the lower three balls and also that the rope has zero thickness, cannot stretch at all, and is not pretensioned, and also that the rope has no friction against the balls.
> Treat this as a statics problem and analyze it. Is it overdetermined? Under what circumstances would the balls move? What is the behavior of the system?
And... first, it says "I’ll separate the geometry from the constraint mechanics, because the key question is not just force balance: it’s whether the inextensible, initially slack-free rope actually fixes the lower-ball geometry or merely limits outward separation." Excuse me? What would the other option be? Either outward separation is prevented or it isn't. Where else could the balls go?
Then, despite the fact that I've mentioned friction in the prompt, it gives an extremely longwinded answer that matches the benchmark and ignores tangential forces entirely without comment. Was it perhaps trained on this crap?
So I followed up:
> Stop ignoring tangential friction forces. I believe that a problem very much like this with a potentially incorrect answer is in your training set. Answer with actual analysis, not based on memory.
Much time was spent thinking. An early part of the answer was "The central correction is this: allowing static friction at the sphere–sphere contacts does not mean arbitrary tangential forces are available. Each sphere must also satisfy torque equilibrium. In this tetrahedral contact geometry, those torque equations force every sphere–sphere tangential contact force to be zero in static equilibrium.". Hey ChatGPT, this is still wrong -- you have forgotten sphere-table friction. The sphere-sphere force on the lower spheres does not have to net out to zero. (And if you do think it nets to zero then you don't need to think any further.)
I then added:
> What if there sphere-table friction?
And encountered the usual problem (which maybe only affects me?) where the chatgpt.com UI becomes kind of unusable after ChatGPT spews too much math into the conversation. But somewhere in the barely-even-scrollable results was this "The three lower balls can be held in position entirely by the coupled sphere–sphere/table friction forces." Hallelujah!
CamperBob2 · · focus · HN ↗
You cannot simply handwave this away. "ChatGPT" models range from GPT 5.5 on "Light" effort to GPT 6 Astra with "Ultra" effort.
The latter yields as good an answer as you could expect for a question that is still poorly formed ("close enough," WTF does that mean?): <a href="https://chatgpt.com/share/6aab1e43-da3c-83e8-9ec4-1b63cec2c141" rel="nofollow">https://chatgpt.com/share/6aab1e43-da3c-83e8-9ec4-1b63cec2c1...
amluto · · focus · HN ↗
Oh, and I don't even have a "light" option, and I'm signed in to a Pro account.
(I continue to despise the chatgpt.com frontend.)
In any case, I tried again forcing GPT-6 Astra Pro (apparently I can't choose GPT-6 Astra non-Pro) and gave the same prompt. It searches the web and gives a long-winded answer, including:
> 7. Does unspecified sphere–sphere friction change the static answer?
> For ordinary ideal point contacts in this regular-tetrahedral configuration, it does not. Gravity, the floor reactions, and the frictionless equatorial rope exert no torque about any sphere’s center. The tangential intersphere forces must therefore balance their torques by themselves.
I'm glad it contemplated the possibility of friction, but it forgot about the table/floor there. So maybe this is a little better than Sol?
[deleted] · · focus · HN ↗
[deleted]
red75prime · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
IanCal · · focus · HN ↗
It didn’t say about prevented vs not, it said about whether the rope fixes them in place (they are all touching) or just limits the separation. Like it’s long enough the balls can be a bit apart but not let the fourth fall fully through.
RomanKornev · · focus · HN ↗
This is definitely not astra, taking a guess this is 5.6, perhaps not even Sol, which does not reflect the state of the frontier (what the research was about).
And yes, the paper is already outdated
amluto · · focus · HN ↗
The paper is about the fact that the benchmark’s evaluator is prone to egregious incorrect rejections of what answers that it should accept. A new model will not invalidate that issue.
emil-lp · · focus · HN ↗
Now, you can't conduct experiments using the default model?
RomanKornev · · focus · HN ↗