‹ BackHN Continuity

Thread

Fable 5 – Median thinking declined in August

428 points · 293 comments · espeed

  1. alexjplant · · focus · HN ↗
    I seem to recall Anthropic going on record saying that they don't do anything to model performance to stretch their compute capacity. I've anecdotally noticed massive peaks and troughs in performance week to week (albeit with Opus, not Fable).

    I wonder what their official explanation for this behavior is.

    1. Wowfunhappy · · focus · HN ↗
      When something is new, its capabilities feel incredible. Over time, those same capabilities become mundane, and you start to notice the flaws.

      (Now, if TFA is actually measuring reasoning tokens, that's quite different! It's not entirely obvious to me how he is measuring.)

      1. knlam · · focus · HN ↗
        Not true. I can read what Fable output with ease but when it sprout Claudish like Opus 5, I know they are doing something to the model. Yes, you can immediate know the claudish language if you work with opus long enough
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.