‹ BackHN Continuity

Thread

Qwen3.8 Max now ranked as the best overall model by agentic index

420 points · 270 comments · apitman

  1. onomojo · · focus · HN ↗
    Any benchmark showing Opus 5 as the best just loses credibility for me. Anyone who's actually used Opus 5 daily knows what I'm talking about.
    1. fellowniusmonk · · focus · HN ↗
      I have some internal tests I use for areas where one particular solution/paradigm is dominant but worse.

      Opus 4.6 is the last model that's actually useful and can "adjust" its perspective to use the newer & better solution.

      Where Opus 4.8-5 has over fit training on worse/older but "dominant" solutions it refuses to adjust.

      Not only does this create an existential threat to adopting progress but it also means that if you have a code base that has rare but real world tradeoff the newest versions of Opus 4.7, 4.8 and 5 are worse than useless and become a major dev timesink.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.