‹ BackHN Continuity

Thread

Qwen3.8 Max now ranked as the best overall model by agentic index

420 points · 270 comments · apitman

  1. onomojo · · focus · HN ↗
    Any benchmark showing Opus 5 as the best just loses credibility for me. Anyone who's actually used Opus 5 daily knows what I'm talking about.
    1. copperx · · focus · HN ↗
      I'm dumbfounded to see Opus 5 making SO MANY mistakes in coding simple stuff. Most times, Fable 5 comes out to be cheaper because it nails so many things much quicker than Opus 5.
      1. TacticalCoder · · focus · HN ↗
        > I'm dumbfounded to see Opus 5 making SO MANY mistakes in coding simple stuff.

        To me it's not so much the dumb mistakes (although there are some of those) but the ultra-verbose, mega-inefficient "solutions" to some problems / prompts.

        Stuff that "works" if you're the kind of person that considers slamming a semi-trailer at 200 mph into a door did, technically, result in the door being somehow "open".

        As it's supposed to be one of the most advanced model, I can't help but wonder if the solutions are that bad/verbose/inefficient because we're already in a loop of models being trained on sloppy-pasta from previous models.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.