‹ BackHN Continuity

Thread

Qwen3.8 Max now ranked as the best overall model by agentic index

403 points · 261 comments · apitman

  1. embedding-shape · · focus · HN ↗
    Strange that the page <a href="https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;agents&#x2F;coding-agents" rel="nofollow">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;agents&#x2F;coding-agents doesn&#x27;t even mention &quot;Qwen&quot; once if it&#x27;s now the &quot;best&quot; according to one of their one index?
    1. scrlk · · focus · HN ↗
      Different benchmarks:

      &gt; Artificial Analysis Agentic Index: Represents the weighted average of agentic capabilities benchmarks in the Artificial Analysis Intelligence Index (GDPval-AA v2, Tau³-Banking)

      &gt; Artificial Analysis Coding Agent Index v1.3 incorporates 3 benchmarks: DeepSWE, Terminal-Bench v2, and SWE-Atlas-QnA

      Qwen3.8 Max is 55.4 on the Agentic Index but hasn&#x27;t been tested for the Coding Agent Index.

      1. apitman · · focus · HN ↗
        Looks like coding agent is model+harness. There are far fewer models represented on that page. I believe &quot;agentic index&quot; is still the metric to look at for coding performance. I could be wrong about that though.
    2. Bootvis · · focus · HN ↗
      Indeed, and this Qwen 3.8 max specific page:

      <a href="https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;qwen3-8-max" rel="nofollow">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;qwen3-8-max

      Doesn&#x27;t have the claim either. Clickbait?

      1. petu · · focus · HN ↗
        This page has it, scroll to &quot;Intelligence&quot; header (not the highlights one, but second on the page &#x2F; with black square) and click &quot;Agentic Index&quot;
        1. Bootvis · · focus · HN ↗
          So the original link should be: <a href="https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;qwen3-8-max?intelligence=agentic-index" rel="nofollow">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;qwen3-8-max?intelligenc...

          Even then, this seems a much more marginal win than the headline suggested to me.

    3. amelius · · focus · HN ↗
      According to those graphs, Grok 4.5 appears to be the most cost-effective model.
      1. user43928 · · focus · HN ↗
        $0.05 per task, Intelligence Index score 52 -&gt; GPT 5.6 Luna max

        $0.36 per task, Intelligence Index score 56 -&gt; Grok 4.5 high

        $1.13 per task, Intelligence Index score 58 -&gt; Qwen 3.8 Max

        $0.81 per task, Intelligence Index score 59 -&gt; GPT 5.6 Sol xhigh

        $1.80 per task, Intelligence Index score 63 -&gt; Opus 5 xhigh

    4. artemisart · · focus · HN ↗
      They didn&#x27;t run all benchmarks. It&#x27;s the best in AA agentic index (GDPval-AA v2, ³-Banking) but not coding index (DeepSWE which is missing, Terminal-Bench v2.1 they have 81% vs 90% for Sol, SWE-Atlas-QnA missing).
    5. moritzwarhier · · focus · HN ↗
      Does &quot;artificial analysis&quot; mean what it says? Dubious.

      But: I&#x27;ve been very impressed by the larger Qwen Models, and a brief try of Kimi also impressed me.

      A lingering sense of quality degradation when going deep remains.

      But that&#x27;s not an accusation: they seem to be hitting the compute&#x2F;quality tradeoff extremely well.

      And on-prem capability is simply irreplaceable.

      Apart from all the innovations that were driven by the strive for this optimization: quantization, &quot;distilling&quot; (without obvious mad-cows-disease)... I think China was an invaluable player in this progress. Intuitively, I&#x27;d even go so far to speculate that LLaMa wouldn&#x27;t exist without the competition.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.