‹ BackHN Continuity

Thread

Qwen3.8 Max now ranked as the best overall model by agentic index

252 points · 127 comments · apitman

  1. embedding-shape · · focus · HN ↗
    Strange that the page <a href="https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;agents&#x2F;coding-agents" rel="nofollow">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;agents&#x2F;coding-agents doesn&#x27;t even mention &quot;Qwen&quot; once if it&#x27;s now the &quot;best&quot; according to one of their one index?
    1. scrlk · · focus · HN ↗
      Different benchmarks:

      &gt; Artificial Analysis Agentic Index: Represents the weighted average of agentic capabilities benchmarks in the Artificial Analysis Intelligence Index (GDPval-AA v2, Tau³-Banking)

      &gt; Artificial Analysis Coding Agent Index v1.3 incorporates 3 benchmarks: DeepSWE, Terminal-Bench v2, and SWE-Atlas-QnA

      Qwen3.8 Max is 55.4 on the Agentic Index but hasn&#x27;t been tested for the Coding Agent Index.

      1. apitman · · focus · HN ↗
        Looks like coding agent is model+harness. There are far fewer models represented on that page. I believe &quot;agentic index&quot; is still the metric to look at for coding performance. I could be wrong about that though.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.