Hi, I’m the author, The Opus 5 substitution was a validation test, not the primary measurement. At that sample size the accuracy difference was -3.8 ± 6.3 points, so it did not clear the pre-registered 99% threshold. Interestingly, output tokens moved much more (-23%), which is why token usage is tracked as a secondary signal.
The actual 10-day windows contain substantially more samples than that validation, but I haven’t demonstrated that they’re sufficient to distinguish a same-family swap of that size, so I’m not claiming they are.
The goal isn’t to make the instrument say “nerfed.” A null result is a result too. I’d much rather publish “we couldn’t detect a change of this magnitude” than overclaim what the data can support.
I have no idea what you're trying to say there about it being a validation test - opus 5 and opus 5.5 are in different universes of ability - if you substituted 5 for 5.5 to validate your system and couldn't tell the difference - your validation failed. Opus 5.5 is replacing a ton of _fable_ usage - if nerfing is real and was of that magnitude there would be no controversy about whether or not it's happening it would be the most obvious thing in the universe
zeroonetwothree · · focus · HN ↗
ninjahawk1 · · focus · HN ↗
The actual 10-day windows contain substantially more samples than that validation, but I haven’t demonstrated that they’re sufficient to distinguish a same-family swap of that size, so I’m not claiming they are.
The goal isn’t to make the instrument say “nerfed.” A null result is a result too. I’d much rather publish “we couldn’t detect a change of this magnitude” than overclaim what the data can support.
vikramkr · · focus · HN ↗