Article: "How Good Are Frontier Models at Physics?
Expert Re-Grading Reveals Broken Evaluations and
Near-Saturation of Leading Benchmarks"
John Sous from Yale posted a fairly solid study on how nearly all the physics benchmarks are broken, as they consistently mark correct answers as incorrect.
When hand grading instead, they found out that the models have actually already saturated the benchmarks which is a little bit scary.
I would be very surprised if any of the frontier models wasn't trained on all public physics benchmarks. Training data providers have been hiring people for exactly this task.
I'd have thought the same but this article from yesterday blew my mind
<a href="https://www.amazon.science/blog/why-dont-machine-learning-research-agents-overfit" rel="nofollow">https://www.amazon.science/blog/why-dont-machine-learning-re...
As it essential implies that these models compress knowledge well in a way that what remains is what's generalizeable more so than remembering every specific detail... Anyway more understanding necessary but thought provoking
qt31415926 · · focus · HN ↗
John Sous from Yale posted a fairly solid study on how nearly all the physics benchmarks are broken, as they consistently mark correct answers as incorrect.
When hand grading instead, they found out that the models have actually already saturated the benchmarks which is a little bit scary.
fsh · · focus · HN ↗
redwood · · focus · HN ↗
As it essential implies that these models compress knowledge well in a way that what remains is what's generalizeable more so than remembering every specific detail... Anyway more understanding necessary but thought provoking