‹ BackHN Continuity

Thread

Qwen3.8 Max now ranked as the best overall model by agentic index

420 points · 270 comments · apitman

  1. onomojo · · focus · HN ↗
    Any benchmark showing Opus 5 as the best just loses credibility for me. Anyone who's actually used Opus 5 daily knows what I'm talking about.
    1. copperx · · focus · HN ↗
      I'm dumbfounded to see Opus 5 making SO MANY mistakes in coding simple stuff. Most times, Fable 5 comes out to be cheaper because it nails so many things much quicker than Opus 5.
      1. usef- · · focus · HN ↗
        Weird how different people's experiences are. If it's making simple mistakes something must be wrong in your setup/context I assume? It's been solid for me, beyond the usual LLMisms that all models have. But I keep context pretty minimal.
        1. efficax · · focus · HN ↗
          Every model that comes out comes with a bunch of people saying "this one is actually dumb they were smart before" and I don't really get it. The models since Opus 4.5 have all been basically the same to me. Sometimes they do the wrong thing, so you have to steer and stop and correct them. Leaving them to operate on their own in no-human-in-the-loop harnesses often gets bad results. But if you single thread it, and keep your work targeted (you have to know what you want the thing to do!), clear your context, the models will do what you ask pretty reliably.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.