‹ BackHN Continuity

Thread

Qwen3.8 Max now ranked as the best overall model by agentic index

420 points · 270 comments · apitman

  1. d2p · · focus · HN ↗
    I clicked through and it showed Qwen at the top at 55.4 compared to 55.3 for Opus Max. I have a screenshot.

    Then I clicked away and back, and now it goes Qwen second, with 58.4, to Opus Max at top with 59.2.

    I have screenshots of both. The description above the chart is the same in boh cases:

    > Artificial Analysis Agentic Index > Represents the weighted average of agentic capabilities benchmarks in the Artificial Analysis Intelligence Index (GDPval-AA v2, ³-Banking)

    What happened? How can the scores change so much in a few seconds?

    1. h14h · · focus · HN ↗
      They JUST updated their methodology:

      <a href="https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;methodology&#x2F;intelligence-benchmarking" rel="nofollow">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;methodology&#x2F;intelligence-bench...

      Edit to provide AA&#x27;s article explaining it:

      <a href="https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;articles&#x2F;artificial-analysis-intelligence-index-v4-1-1" rel="nofollow">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;articles&#x2F;artificial-analysis-i...

      1. splatzone · · focus · HN ↗
        Can someone please explain what changed, when it happened, and whether it was surreptitious?
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.