‹ BackHN Continuity

Thread

Qwen3.8 Max now ranked as the best overall model by agentic index

252 points · 127 comments · apitman

  1. SwellJoe · · focus · HN ↗
    I find that surprising.

    I've been trying it on several projects and have found it's pretty sloppy. It leaves stuff broken, doesn't reliably write tests to check its own work unless explicitly prompted, misunderstands the assignment, etc.

    It is smart and reasonably quick but not reliable.

    1. dyauspitr · · focus · HN ↗
      It’s because they’re doing some sort of combined score of intelligence, speed and cost. On pure intelligence it doesn’t even show up in the top 10.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.