‹ BackHN Continuity

Thread

Qwen3.8 Max now ranked as the best overall model by agentic index

403 points · 261 comments · apitman

  1. onomojo · · focus · HN ↗
    Any benchmark showing Opus 5 as the best just loses credibility for me. Anyone who's actually used Opus 5 daily knows what I'm talking about.
    1. enraged_camel · · focus · HN ↗
      It's my daily driver. I like it and find it noticeably better than Opus 4.8.

      After I started reading complaints about Opus 5, I gave Fable the task of evaluating a bunch of code Opus 4.8 had written and compare it to Opus 5's code. Fable ran a dynamic workflow and the scores came back 15-20% higher for Opus 5's code in terms of quality, correctness and readability/conciseness. I did not tell Fable which Opus wrote which code, and I turned off memory as well to ensure there was no pollution from that angle.

      My only complaint is that Opus 5's prose is annoying as hell. I wrote a custom skill for it for concise debriefs and it has been working pretty well for me.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.