They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They're currently tracking Opus 5.5 and GPT-6 Astra.
This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it's often about honeymoon effects.
Anthropic A/Bs my weekly quota amount. So I have an automated prompt that runs at 3 AM with a transcription task, I measure input and output tokens, and weekly/5 hour quota before and after. The absolute token counts stay within 0.1% while in mode A it counts for 1% of my 5 hour quota and mode B 4% of my 5 hour quota.
Do Anthropic quotas give you precise remaining token counts or something? I have something similar set up for tracking my ChatGPT usage but it only gives percentages remaining, which is a pretty coarse metric.
Claude code supposedly has otel you can set via env. I haven't set it up, so I'm just repeating hearsay.. but it supposedly has everything relevant in it wrt token usage and cost
It's meant for their test env I think, so is not documented to my knowledge
It's for corporate users who want to track how the product is used internally, and it was documented at least some time ago, quite extensively even.
It's obvious unless you have multiple requests from different models in flight at the same time, and the sum total usage comes out to less than a single percentage.
jug · · focus · HN ↗
<a href="https://www.bridgebench.ai/nerf-bench" rel="nofollow">https://www.bridgebench.ai/nerf-bench
They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They're currently tracking Opus 5.5 and GPT-6 Astra.
This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it's often about honeymoon effects.
user3939382 · · focus · HN ↗
apitman · · focus · HN ↗
ffsm8 · · focus · HN ↗
It's meant for their test env I think, so is not documented to my knowledge
adastra22 · · focus · HN ↗
reubenmorais · · focus · HN ↗
TeMPOraL · · focus · HN ↗
ffsm8 · · focus · HN ↗
<a href="https://code.claude.com/docs/en/monitoring-usage#usage-monitoring" rel="nofollow">https://code.claude.com/docs/en/monitoring-usage#usage-monit...
thanks for correcting me on that regard
Aeolun · · focus · HN ↗
apitman · · focus · HN ↗