‹ BackHN Continuity

Thread

Fable 5 – Median thinking declined in August

428 points · 293 comments · espeed

  1. alexjplant · · focus · HN ↗
    I seem to recall Anthropic going on record saying that they don't do anything to model performance to stretch their compute capacity. I've anecdotally noticed massive peaks and troughs in performance week to week (albeit with Opus, not Fable).

    I wonder what their official explanation for this behavior is.

    1. Wowfunhappy · · focus · HN ↗
      When something is new, its capabilities feel incredible. Over time, those same capabilities become mundane, and you start to notice the flaws.

      (Now, if TFA is actually measuring reasoning tokens, that's quite different! It's not entirely obvious to me how he is measuring.)

      1. chrsw · · focus · HN ↗
        I don’t think that’s what’s going on. I notice flaws on day one of model releases. But I also notice improvements if the model is truly more advanced than what I’m used to. Then over time the same questions or tasks return worse results.

        What is actually stopping these model companies from running a model at full capacity on release then once its name rings out, start serving users quantized garbage?

        1. dist-epoch · · focus · HN ↗
          It's called hedonic adaptation.

          > What is actually stopping these model companies

          You can say this about any company in the world, selling anything.

          It's trivially measurable, and there are people running the same benchmark on the leading models every day and measuring if they degrade. Spoiler: they don't.

          But you can always say "the conspiracy goes higher", and that the companies know about these daily benchmarks and are routing them to "quality" envs.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.