‹ BackHN Continuity

Thread

Qwen3.8 Max now ranked as the best overall model by agentic index

403 points · 261 comments · apitman

  1. d2p · · focus · HN ↗
    I clicked through and it showed Qwen at the top at 55.4 compared to 55.3 for Opus Max. I have a screenshot.

    Then I clicked away and back, and now it goes Qwen second, with 58.4, to Opus Max at top with 59.2.

    I have screenshots of both. The description above the chart is the same in boh cases:

    > Artificial Analysis Agentic Index > Represents the weighted average of agentic capabilities benchmarks in the Artificial Analysis Intelligence Index (GDPval-AA v2, ³-Banking)

    What happened? How can the scores change so much in a few seconds?

    1. h14h · · focus · HN ↗
      They JUST updated their methodology:

      <a href="https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;methodology&#x2F;intelligence-benchmarking" rel="nofollow">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;methodology&#x2F;intelligence-bench...

      Edit to provide AA&#x27;s article explaining it:

      <a href="https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;articles&#x2F;artificial-analysis-intelligence-index-v4-1-1" rel="nofollow">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;articles&#x2F;artificial-analysis-i...

      1. ahartmetz · · focus · HN ↗
        Fixed the result, eh? In both senses of the word.
      2. gpt5 · · focus · HN ↗
        What was the change?
      3. johnnyApplePRNG · · focus · HN ↗
        I have been suspicious of these AI leaderboard sites for some time now, and this only affirms that suspicion.
      4. torginus · · focus · HN ↗
        In that case they should clearly label that this is a new benchmark.
      5. splatzone · · focus · HN ↗
        Can someone please explain what changed, when it happened, and whether it was surreptitious?
      6. kmeh · · focus · HN ↗
        &gt; HLE, AA-LCR and AA-Omniscience are now graded by GPT-5.6 Luna (medium), replacing GPT-4o, Qwen3 235B A22B 2507, and Gemini 3 Flash Preview respectively. These checks are now unified under a more capable modern model, selected for strong agreement with human judgment in our grader validation

        Interesting that they chose a nano-sized model from OpenAI to be a grader for benchmarks involving knowledge and hallucination.

        1. nolok · · focus · HN ↗
          What&#x27;s interesting is that if you ask 5.6 Sol or Opus 5 they will tell you it&#x27;s a bad idea to have the reviewer be the dumber of the set as it can&#x27;t judge them properly to decide who is right, and thus if one is better because it found an answer that&#x27;s better but contradict the obvious it would be biased against. I know because I just had a consensus conversation with them this afternoon about a design that was similar (though about something completly different than judging agentic quality or whatever).
    2. WD-42 · · focus · HN ↗
      Same, they just updated it. Hacker news effect?
    3. apitman · · focus · HN ↗
      Welp
    4. personjerry · · focus · HN ↗
      They should probably freeze the results before publishing.
    5. Gcam · · focus · HN ↗
      Hey! George from the Artificial Analysis team here. We published an update today that does result in a change of the order, Qwen3.8 Max to second rather than first. The methodology change was an already planned upgrade to our equality checking&#x2F;grader models, and brings the latest ³-Banking version to Artificial Analysis. Regular updates are normal for us to keep our benchmarks up to date.

      The order changes but I think the story discussed in this thread holds - this is a very impressive release and Qwen3.8 Max is a huge step up in agentic capabilities.

      Relevant blog post (also linked to by others): <a href="https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;articles&#x2F;artificial-analysis-intelligence-index-v4-1-1" rel="nofollow">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;articles&#x2F;artificial-analysis-i...

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.