‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

741 points · 330 comments · snehesht

  1. SuperV1234 · · focus · HN ↗
    We're getting closer and closer to the day we can have an Opus-like model running locally. The dream!
    1. stymaar · · focus · HN ↗
      It depends on what you mean by “Opus-like”, because if you mean “as strong as Opus 4.6 for agentic coding” then Qwen3.8-27B has been there for the past two months.

      But if you mean “as strong as current-gen Opus” then it's probably never gonna happen, but it doesn't really matter since we're long into the diminishing returns for performance improvements: I haven't notice any major leap between 4.6 and 5.5 in my daily usage, and I'm convinced that with a fact enough piecs of hardware I would be using local Qwen exclusively (I'm using it daily but only at night for long running tasks because they take much more time than Opus due to the compounding effects of my slow GPU and Qwen's verbosity).

      1. nycdatasci · · focus · HN ↗
        Opus 5.5 was a step function change. Really crushes on multi-hour coding compared with prior models.
        1. jjcm · · focus · HN ↗
          Highly agree.

          I keep saying “I’d be so Happy with ${currentOpusVersion} locally”, but I keep being impressed with how much the capabilities change between versions. I have a RTX 6000 pro so I can easily run this qwen 3.8 flash next, but it’s much harder to give up the freedom that 5.5 gives me.

        2. stymaar · · focus · HN ↗
          I feel like Anthropic is going back and forth on this, as Opus 5 was the laziest since 4.6 by far, often giving up after ten or fifteen minutes and being like “of course there's still this, and this and this to do, but it's long and you should take time to think about the schedule for these tasks”. It gave up way earlier than Qwen3.8-27B for instance.

          Opus 5.5 just fell like a return to the mean afterwards (it's probably an improvement over the previous versions, but I couldn't really feel it because I've been burned by Opus 5).

      2. hgoel · · focus · HN ↗
        The most visible leap between 4.6 and 5.5 seems to be that the latter has gotten much more computer-use training, so there's a clear progression in the ability for the model to use Blender. But catching up on that is just a matter of training on the same thing.
      3. system2 · · focus · HN ↗
        I would be forever happy with Opus 4.8.
        1. bitexploder · · focus · HN ↗
          Flash next is /really/ close. It is at parity with 4.7 as far as I can tell and basically where Opus 4.8 was. It is a genuinely good model. And I run it at home on $1500 of GPU at 125 t/s :)
          1. gruturo · · focus · HN ↗
            Seconding this. Flash-next (and let's not forget, it's a PREVIEW of the 4 architecture - with the "real" 4 rumored coming later this month) is the first model I can run on reasonable hardware (2 thoroughly obsolete P100s off ebay at ~$100 each plus the RAM I could scavenge from other PCs at home) at a reasonable speed (22, with GPUs in layer-split due to llama-cpp's limitation on qwen4-exp arch, and no MTP. Strata could double these numbers).

            It's.... the real thing, for the first time. If you cut me off cloud models today, I would get plenty of utility out of this thing.

            (Others may have had the same feeling from GLM5.3 or Deepseek 4.1 flash but I never had a chance of running those.)

          2. fsiefken · · focus · HN ↗
            Running an nvidia card at full load, would cost me ~100 euro of electricty each month (europe). Of course one wouldn't have usage caps.

            Where the internet was a subscription 15 euro subscription to encyclopaedic knowledge, an genAI subscription is renting a researcher/programmer for 100 euro.

            1. stymaar · · focus · HN ↗
              > (europe)

              Given the massive difference in electricity price between different european countries, adding "Europe" doesn't bring much context.

              1. magicalhippo · · focus · HN ↗
                Even with a country. Here in Norway we have multiple price zones, and at times there can be 100x difference between them, often 10x. All due to lack of transmission capacity between northern and southern zones.
                1. stymaar · · focus · HN ↗
                  TIL, thanks.

                  In France we have the same price in all continental France, there are whole regions that are poorly connected to the grid (Britanny and the French Riviera) or even remote islands like Corsica or even the islands in the Indian Ocean and the West Indies but the price is still the same as elsewhere.

            2. bitexploder · · focus · HN ↗
              These GPUs use about 400W at full tilt (200W each), for reference.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.