‹ BackHN Continuity

Thread

Qwen3.8 Max now ranked as the best overall model by agentic index

292 points · 173 comments · apitman

  1. d2p · · focus · HN ↗
    I clicked through and it showed Qwen at the top at 55.4 compared to 55.3 for Opus Max. I have a screenshot.

    Then I clicked away and back, and now it goes Qwen second, with 58.4, to Opus Max at top with 59.2.

    I have screenshots of both. The description above the chart is the same in boh cases:

    > Artificial Analysis Agentic Index > Represents the weighted average of agentic capabilities benchmarks in the Artificial Analysis Intelligence Index (GDPval-AA v2, ³-Banking)

    What happened? How can the scores change so much in a few seconds?

    1. h14h · · focus · HN ↗
      They JUST updated their methodology:

      <a href="https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;methodology&#x2F;intelligence-benchmarking" rel="nofollow">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;methodology&#x2F;intelligence-bench...

      1. ahartmetz · · focus · HN ↗
        Fixed the result, eh? In both senses of the word.
      2. gpt5 · · focus · HN ↗
        What was the change?
      3. johnnyApplePRNG · · focus · HN ↗
        I have been suspicious of these AI leaderboard sites for some time now, and this only affirms that suspicion.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.