‹ BackHN Continuity

Thread

Qwen3.8 Max now ranked as the best overall model by agentic index

420 points · 270 comments · apitman

  1. d2p · · focus · HN ↗
    I clicked through and it showed Qwen at the top at 55.4 compared to 55.3 for Opus Max. I have a screenshot.

    Then I clicked away and back, and now it goes Qwen second, with 58.4, to Opus Max at top with 59.2.

    I have screenshots of both. The description above the chart is the same in boh cases:

    > Artificial Analysis Agentic Index > Represents the weighted average of agentic capabilities benchmarks in the Artificial Analysis Intelligence Index (GDPval-AA v2, ³-Banking)

    What happened? How can the scores change so much in a few seconds?

    1. h14h · · focus · HN ↗
      They JUST updated their methodology:

      <a href="https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;methodology&#x2F;intelligence-benchmarking" rel="nofollow">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;methodology&#x2F;intelligence-bench...

      Edit to provide AA&#x27;s article explaining it:

      <a href="https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;articles&#x2F;artificial-analysis-intelligence-index-v4-1-1" rel="nofollow">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;articles&#x2F;artificial-analysis-i...

      1. kmeh · · focus · HN ↗
        &gt; HLE, AA-LCR and AA-Omniscience are now graded by GPT-5.6 Luna (medium), replacing GPT-4o, Qwen3 235B A22B 2507, and Gemini 3 Flash Preview respectively. These checks are now unified under a more capable modern model, selected for strong agreement with human judgment in our grader validation

        Interesting that they chose a nano-sized model from OpenAI to be a grader for benchmarks involving knowledge and hallucination.

        1. nolok · · focus · HN ↗
          What&#x27;s interesting is that if you ask 5.6 Sol or Opus 5 they will tell you it&#x27;s a bad idea to have the reviewer be the dumber of the set as it can&#x27;t judge them properly to decide who is right, and thus if one is better because it found an answer that&#x27;s better but contradict the obvious it would be biased against. I know because I just had a consensus conversation with them this afternoon about a design that was similar (though about something completly different than judging agentic quality or whatever).
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.