The Office of Weights and Measures exists because, long before any of us were born, in 1836, companies were up to shady shit and consumers were paying for inconsistent products. I.E. Being scammed.
AI companies should be subject to the OWM like any other company that sells a product that varies in weight. Perhaps when a sane administration is re-elected; one that can read history books and comprehend why our regulations exist in the first place. Or have even a semblance of respect for its citizenry.
I doubt that would change the perception. Every model release is followed by accusations of nerfing.
There are several projects that repeat benchmarks on published models. None has ever found significant fluctations
Here's one example <a href="https://marginlab.ai/trackers/claude-code/" rel="nofollow">https://marginlab.ai/trackers/claude-code/
Fluctuations of a few percentage points are to be expected and should not surprise anyone who knows how LLMs work.
This Twitter analysis of Fable 5 is not that at all. They analyzed their coding sessions and blamed all of the fluctuations on Fable changing. They then compared to ARC-AGI-2 questions as the benchmark for thinking tokens and tried to stir up anger that coding turns don't produce as many thinking tokens as the ARC-AGI-2 problems.
If you go to <a href="https://marginlab.ai/trackers/claude-code-historical-performance/" rel="nofollow">https://marginlab.ai/trackers/claude-code-historical-perform... there is a very clear downwards trend in the two weeks before Opus 4.7 release. Then a sudden and dramatic drop seven days before Opus 4.8. And now we seem to have entered another decline in the last ten days, beyond the usual noise of Opus 5 scores
jesse_dot_id · · focus · HN ↗
AI companies should be subject to the OWM like any other company that sells a product that varies in weight. Perhaps when a sane administration is re-elected; one that can read history books and comprehend why our regulations exist in the first place. Or have even a semblance of respect for its citizenry.
Aurornis · · focus · HN ↗
There are several projects that repeat benchmarks on published models. None has ever found significant fluctations
Here's one example <a href="https://marginlab.ai/trackers/claude-code/" rel="nofollow">https://marginlab.ai/trackers/claude-code/
Fluctuations of a few percentage points are to be expected and should not surprise anyone who knows how LLMs work.
This Twitter analysis of Fable 5 is not that at all. They analyzed their coding sessions and blamed all of the fluctuations on Fable changing. They then compared to ARC-AGI-2 questions as the benchmark for thinking tokens and tried to stir up anger that coding turns don't produce as many thinking tokens as the ARC-AGI-2 problems.
wongarsu · · focus · HN ↗