‹ BackHN Continuity

Thread

Qwen3.8 Max now ranked as the best overall model by agentic index

449 points · 287 comments · apitman

  1. d2p · · focus · HN ↗
    I clicked through and it showed Qwen at the top at 55.4 compared to 55.3 for Opus Max. I have a screenshot.

    Then I clicked away and back, and now it goes Qwen second, with 58.4, to Opus Max at top with 59.2.

    I have screenshots of both. The description above the chart is the same in boh cases:

    > Artificial Analysis Agentic Index > Represents the weighted average of agentic capabilities benchmarks in the Artificial Analysis Intelligence Index (GDPval-AA v2, ³-Banking)

    What happened? How can the scores change so much in a few seconds?

    1. Gcam · · focus · HN ↗
      Hey! George from the Artificial Analysis team here. We published an update today that does result in a change of the order, Qwen3.8 Max to second rather than first. The methodology change was an already planned upgrade to our equality checking/grader models, and brings the latest ³-Banking version to Artificial Analysis. Regular updates are normal for us to keep our benchmarks up to date.

      The order changes but I think the story discussed in this thread holds - this is a very impressive release and Qwen3.8 Max is a huge step up in agentic capabilities.

      Relevant blog post (also linked to by others): <a href="https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;articles&#x2F;artificial-analysis-intelligence-index-v4-1-1" rel="nofollow">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;articles&#x2F;artificial-analysis-i...

      1. saretup · · focus · HN ↗
        You gotta admit the timing looks very suspicious.
        1. TacticalCoder · · focus · HN ↗
          &gt; You gotta admit the timing looks very suspicious.

          Do you mean the timing looks like: &quot;We&#x27;re SV tech-bros. Our benchmarks showed a chinese model above what&#x27;s considered the best model at the moment. So we quickly modified the benchmark so that our SV tech-bros don&#x27;t look like they&#x27;re losing to a chinese model&quot;?

          That&#x27;s indeed a bit fishy.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.