‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

819 points · 362 comments · snehesht

  1. snehesht · · focus · HN ↗
    I tried it and it worked surprisingly well. On my machine (Nvidia 4090, 128GB DDR5, Ryzen 7950x3d) I'm getting 124 tokens per sec, thought to share it here.

    <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;Qwen&#x2F;Qwen3.8-Flash-Next" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;Qwen&#x2F;Qwen3.8-Flash-Next

    1. roscas · · focus · HN ↗
      Coder version with 30t&#x2F;sec on a Ryzen 3600x with 48GB of RAM with a nvidia 3080.

      This is not a very fast desktop. Memory speed is around 2000mhz only. My SSD is some of the worst SSD I&#x27;ve seen and 3080 had its days of glory.

      I still have code, chromium, librewolf and many other programs running. I have video streams running while I also watch tv and many times youtube videos.

      I use it with the browser that has a great dashboard and with hermes agent and that it really makes this amazing.Only change I made is to set thinking to low.

      This is a coding model. Any other task, I still use Ornith 1.5 35B that throws 20t&#x2F;sec and Laguna.XS-2.0.

      1. StumpChunkman · · focus · HN ↗
        How much VRAM on your 3080? I&#x27;ve got an early 10gb model. I&#x27;ve been thinking of exploring local coding models, but everyone seems to use much better GPUs than I have access to. Yours is one of the first I&#x27;ve seen with maybe similar hardware on some level.
        1. roscas · · focus · HN ↗
          Yes, 3080 with 10GB, forgot to mention that.

          Mine is at the moment writting some cpp code for some SBOM tests.

          I have loads of terminals open. Librewolf, Chromium and you know how this crap likes ram, I have also a vm with 4gb of ram running and doing stuff while I wait for the results but hey, while I wrote this the program is done. Wow! That was 29.x tokens per second most of the time.

          Oh I will run some other tests with hermes now because hermes is amazing too.

        2. Abishek_Muthian · · focus · HN ↗
          With limited VRAM, you can still use local LLMs but you have to set your expectations correctly to solve the right problem.

          I run small models on various kind of devices including 1.7B model on the original Jetson Nano (4B) abandoned by Nvidia, I had to upscale the software (OS, Lllama.cpp etc.) to run the model but the model runs at 17t&#x2F;s.

          Small models are great at NLP stuff like classification (e.g. bookmarking), ASR etc. I have built custom browser extensions to save time with bookmarking and categorizing to use with local LLMs.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.