‹ BackHN Continuity

Thread

Qwen3.8 Max now ranked as the best overall model by agentic index

403 points · 261 comments · apitman

  1. onomojo · · focus · HN ↗
    Any benchmark showing Opus 5 as the best just loses credibility for me. Anyone who's actually used Opus 5 daily knows what I'm talking about.
    1. cromka · · focus · HN ↗
      Agreed, it's extremely frustrating. It's the only model that actually makes me curse when talking to it, even knowing how counterproductive it is.
      1. moffkalast · · focus · HN ↗
        I'd certainly rank it at the very top of the want to kill yourself when using it benchmark. It outperforms everything else on that leaderboard.

        With weaker models you can sort of understand, they're trying their best and failing, but this thing just channels its immense inteligence into being as annoying as possible instead. I know it can do what I'm asking it to do, but it just finds a way to weasel out of it, or maybe just thinks for 10 minutes instead, then fixes one thing and breaks four additional ones.

        1. msp26 · · focus · HN ↗
          yep matches my experience completely

          But even fable has the annoying tendency to invent new jargon and produce an incomprehensible soup of text.

          1. cromka · · focus · HN ↗
            Is there any model that knows how to smooth an overly literary text over? I find Opus and Fable constantly decorate the documentation they write like a damn 19/20th century writer. We're working with IT stuff yet it writes like it's going to win some Pulitzer prize.
            1. msp26 · · focus · HN ↗
              Not sure how to fully fix this but I remember a session last week where I got so fed up mid way though reading a response that I used the following:

              "give me this again without jargon invented this session at high density

              and with a couple (maybe more or less) simple useful ascii diagrams underneath each design"

              The context is that I was discussing an experimental new idea for my video game review analysis product.

              Designs 1 and 2 were horrible: the model even suggested a rejection after the word soup so it would have been pointless to waste my fleeting life on earth reading it.

              Otherwise, I generally really enjoyed using fable for bouncing ideas. It was an absolute joy to have this thing provide useful criticism, analyse sample data, and create prototypes so that I could elevate my understanding of the problem without stepping down from a pure intuition/design headspace.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.