‹ BackHN Continuity

Thread

Fable 5 – Median thinking declined in August

428 points · 293 comments · espeed

  1. Waterluvian · · focus · HN ↗
    I have no hard data but I have a strong feeling this morning that something's wrong with Fable 5 compared to Friday evening.

    Just an hour ago I had Fable correctly identify an unused method that could be deleted. I then immediately get a diff for an exact duplicate method, and then Fable outputting, "I accidentally duplicated <method> instead of deleting it. Removing both copies now."

    The remaining morning complaints that makes it feel like something's off is that it will do a lot of "thinking" for simple things that previously took very little time. And it got very lost and completely mixed up DE-91M predicate names and implementations. Just absolute disaster code that I had over the past months come to generally expect it to do without issue.

    Glad I carefully review everything. I think what I need is reliability and consistency. But it feels like picking a model from the list doesn't guarantee that: that the models' "brain" is open on the table and they're screwing with it.

    1. zarmin · · focus · HN ↗
      I would rather wait in a queue than be routed to a degraded model. And if they _have_ to degrade the models, then I wish they would fucking tell us. Instead, it's "I have a strong feeling".

      That we have to guess at this is by far the worst part of the AI era. It feels like a dark cloud over my productivity. It makes my body tense for the entire day when it happens. Not healthy.

      1. meowface · · focus · HN ↗
        They have repeatedly said they do not ever intentionally reduce model quality and do not degrade in this way, and that a model version number is always the same.

        But, of course, OP is an empirical claim to the contrary, and I'd be curious to see if anyone (who's been capturing data over these timeframes) can replicate the same results and if Anthropic has any comment.

        1. mh- · · focus · HN ↗
          Every official statement I've seen around this is careful to say that they "don't intentionally reduce model quality", which leaves plenty of room for "we adjusted some knobs and our evals show performance is materially the same".

          However, I also agree that I haven't seen any robust data from someone tracking it daily/weekly. The handful of sites purporting to do this aren't even running it enough times to hit stat sig.

          edit: someone linked one elsewhere in this thread called AI Stupid Level - they "run 7 trials instead of just 1". I don't blame them. Doing this in a statistically sound manner would cost a small fortune.

          1. meowface · · focus · HN ↗
            I kind of feel "reduce the amount of thinking tokens produced" would fall under degrading model quality.

            In any case, I am willing to believe it's possible something degraded, but so far I have not seen any empirical evidence of it since the previous incident with the inference and harness bugs. I lean towards Anthropic probably not intentionally doing anything like this without disclosing it beforehand.

            1. pixl97 · · focus · HN ↗
              The issue here is you have to think like a lawyer trying to weasel out of making an empirical statement.

              For example "We didn't change any settings, but when GPU use gets high the run time of a prompt is lessened. But you must remember this is always in effect so nothing changed at all. This happens occasionally on random prompts some of the time, and when it's busy it happens all of the time".

              In someones eye this would fit the letter of the law but not the spirit of the law that you hold.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.