‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

741 points · 330 comments · snehesht

  1. snehesht · · focus · HN ↗
    I tried it and it worked surprisingly well. On my machine (Nvidia 4090, 128GB DDR5, Ryzen 7950x3d) I'm getting 124 tokens per sec, thought to share it here.

    <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;Qwen&#x2F;Qwen3.8-Flash-Next" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;Qwen&#x2F;Qwen3.8-Flash-Next

    1. proc0 · · focus · HN ↗
      Do you know how it compares to Qwen 3.8 27B? I really want to compare the distilled ones with harness versus the full MoE versions.
      1. incognito124 · · focus · HN ↗
        Qwen 3.8 flash next is way better than 27B. It&#x27;s so good I dont even use claude anymore
        1. snehesht · · focus · HN ↗
          Yeah I agree, I&#x27;m running it with Pi didn&#x27;t notice much difference compared to lower tier models and the speed, of course.
          1. nicce · · focus · HN ↗
            I am running 27B with Deepseek Harness these days and somehow just by using it, without any parameter changes, the model feels even more intelligent.
            1. DoctorOetker · · focus · HN ↗
              do LLMs tend to be homesick when not used in the same harness they sat in during some training phase?
              1. Bnjoroge · · focus · HN ↗
                iirc there was a sectionin Qwen’s paper where they talked anout how they post-trained flash or 3.8 to work just as well regardless of the harness or eval used. I think that used to be true but not sure if it is any longer
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.