‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

741 points · 330 comments · snehesht

  1. XCSme · · focus · HN ↗
    For what's worth, on high, 27b is better, but both models on high are too slow to be used (way too many output tokens).

    On low, 3.8 Flash Next is better AND considerably faster.

    My results[0][1], on a RTX 3090 + 128GB DDR4 RAM. I still have to see if I can optimize any more settings for either, but I think I will switch from 3.8 27b to 3.8 Flash Next for my local LLM uses.

    [0]: <a href="https:&#x2F;&#x2F;aibenchy.com&#x2F;?q=qwen+3.8+27b%2C+qwen+3.8+next" rel="nofollow">https:&#x2F;&#x2F;aibenchy.com&#x2F;?q=qwen+3.8+27b%2C+qwen+3.8+next

    [1]: <a href="https:&#x2F;&#x2F;aibenchy.com&#x2F;compare&#x2F;qwen-qwen3-8-27b-low&#x2F;qwen-qwen3-8-flash-next-low&#x2F;qwen-qwen3-8-27b-high&#x2F;qwen-qwen3-8-flash-next-xhigh&#x2F;" rel="nofollow">https:&#x2F;&#x2F;aibenchy.com&#x2F;compare&#x2F;qwen-qwen3-8-27b-low&#x2F;qwen-qwen3...

    1. kelvie · · focus · HN ↗
      These benchmarks seems to suggest that Flash Next performs best on Low thinking compared to the same model on higher thinking levels?

      And that Qwen 27b outperforms both when set to high?

      What quants are being compared here?

      1. XCSme · · focus · HN ↗
        Both models on high are kinda bugged and don&#x27;t give better results. Use them on low, unless you are generating images&#x2F;art&#x2F;design.

        The exact quant is mentioned on model page[0] IQ3_S, and I think 27b was Q4, via Ollama, the one that fits on a 3090 24GB

        [0]: <a href="https:&#x2F;&#x2F;aibenchy.com&#x2F;model&#x2F;qwen-qwen3-8-flash-next-low&#x2F;" rel="nofollow">https:&#x2F;&#x2F;aibenchy.com&#x2F;model&#x2F;qwen-qwen3-8-flash-next-low&#x2F;

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.