I've followed a few trackers, eg <a href="https://marginlab.ai/trackers/claude-code/" rel="nofollow">https://marginlab.ai/trackers/claude-code/ , for awhile. For Claude Code the trend, it seems to me at least, is fewer tokens to do the same or better job. Prompt changes, tool ergonomics changes, etc.; I'd be shocked if they didn't A/B every release. Less thinking as measured by tokens isn't necessarily bad if you can get the same results by making it think about the "right" things or structure. They obviously screw up sometimes, and I've always been suspicious with hidden tokens, but I haven't found evidence quality intentionally degrades over time.
These analyses are much better than these Twitter charts.
I don't think anyone is reading the details for the Twitter post because it was not an actual benchmark. They did a post-hoc analysis of their logs from day to day.
Their random collection of prompts for each day is not a benchmark.
The site you linked is a much better example of a real benchmark being repeated over time.
rcr-anti · · focus · HN ↗
Aurornis · · focus · HN ↗
I don't think anyone is reading the details for the Twitter post because it was not an actual benchmark. They did a post-hoc analysis of their logs from day to day.
Their random collection of prompts for each day is not a benchmark.
The site you linked is a much better example of a real benchmark being repeated over time.