‹ BackHN Continuity

Thread

Fable 5 – Median thinking declined in August

428 points · 293 comments · espeed

  1. rcr-anti · · focus · HN ↗
    I&#x27;ve followed a few trackers, eg <a href="https:&#x2F;&#x2F;marginlab.ai&#x2F;trackers&#x2F;claude-code&#x2F;" rel="nofollow">https:&#x2F;&#x2F;marginlab.ai&#x2F;trackers&#x2F;claude-code&#x2F; , for awhile. For Claude Code the trend, it seems to me at least, is fewer tokens to do the same or better job. Prompt changes, tool ergonomics changes, etc.; I&#x27;d be shocked if they didn&#x27;t A&#x2F;B every release. Less thinking as measured by tokens isn&#x27;t necessarily bad if you can get the same results by making it think about the &quot;right&quot; things or structure. They obviously screw up sometimes, and I&#x27;ve always been suspicious with hidden tokens, but I haven&#x27;t found evidence quality intentionally degrades over time.
    1. Aurornis · · focus · HN ↗
      These analyses are much better than these Twitter charts.

      I don&#x27;t think anyone is reading the details for the Twitter post because it was not an actual benchmark. They did a post-hoc analysis of their logs from day to day.

      Their random collection of prompts for each day is not a benchmark.

      The site you linked is a much better example of a real benchmark being repeated over time.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.