Qwen3.8 Max now ranked as the best overall model by agentic index
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Qwen3.8 Max now ranked as the best overall model by agentic index
Unofficial Hacker News client; not affiliated with Y Combinator.
embedding-shape · · focus · HN ↗
scrlk · · focus · HN ↗
> Artificial Analysis Agentic Index: Represents the weighted average of agentic capabilities benchmarks in the Artificial Analysis Intelligence Index (GDPval-AA v2, Tau³-Banking)
> Artificial Analysis Coding Agent Index v1.3 incorporates 3 benchmarks: DeepSWE, Terminal-Bench v2, and SWE-Atlas-QnA
Qwen3.8 Max is 55.4 on the Agentic Index but hasn't been tested for the Coding Agent Index.
apitman · · focus · HN ↗
Bootvis · · focus · HN ↗
<a href="https://artificialanalysis.ai/models/qwen3-8-max" rel="nofollow">https://artificialanalysis.ai/models/qwen3-8-max
Doesn't have the claim either. Clickbait?
petu · · focus · HN ↗
Bootvis · · focus · HN ↗
Even then, this seems a much more marginal win than the headline suggested to me.
amelius · · focus · HN ↗
user43928 · · focus · HN ↗
$0.36 per task, Intelligence Index score 56 -> Grok 4.5 high
$1.13 per task, Intelligence Index score 58 -> Qwen 3.8 Max
$0.81 per task, Intelligence Index score 59 -> GPT 5.6 Sol xhigh
$1.80 per task, Intelligence Index score 63 -> Opus 5 xhigh
artemisart · · focus · HN ↗
moritzwarhier · · focus · HN ↗
But: I've been very impressed by the larger Qwen Models, and a brief try of Kimi also impressed me.
A lingering sense of quality degradation when going deep remains.
But that's not an accusation: they seem to be hitting the compute/quality tradeoff extremely well.
And on-prem capability is simply irreplaceable.
Apart from all the innovations that were driven by the strive for this optimization: quantization, "distilling" (without obvious mad-cows-disease)... I think China was an invaluable player in this progress. Intuitively, I'd even go so far to speculate that LLaMa wouldn't exist without the competition.