‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

741 points · 330 comments · snehesht

  1. SuperV1234 · · focus · HN ↗
    We're getting closer and closer to the day we can have an Opus-like model running locally. The dream!
    1. stymaar · · focus · HN ↗
      It depends on what you mean by “Opus-like”, because if you mean “as strong as Opus 4.6 for agentic coding” then Qwen3.8-27B has been there for the past two months.

      But if you mean “as strong as current-gen Opus” then it's probably never gonna happen, but it doesn't really matter since we're long into the diminishing returns for performance improvements: I haven't notice any major leap between 4.6 and 5.5 in my daily usage, and I'm convinced that with a fact enough piecs of hardware I would be using local Qwen exclusively (I'm using it daily but only at night for long running tasks because they take much more time than Opus due to the compounding effects of my slow GPU and Qwen's verbosity).

      1. nycdatasci · · focus · HN ↗
        Opus 5.5 was a step function change. Really crushes on multi-hour coding compared with prior models.
        1. stymaar · · focus · HN ↗
          I feel like Anthropic is going back and forth on this, as Opus 5 was the laziest since 4.6 by far, often giving up after ten or fifteen minutes and being like “of course there's still this, and this and this to do, but it's long and you should take time to think about the schedule for these tasks”. It gave up way earlier than Qwen3.8-27B for instance.

          Opus 5.5 just fell like a return to the mean afterwards (it's probably an improvement over the previous versions, but I couldn't really feel it because I've been burned by Opus 5).

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.