‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

806 points · 356 comments · snehesht

  1. esafak · · focus · HN ↗
    Has anyone calculated the effective intelligence of these quantized models?

    I think publishing benchmarks with quantized models should become standard practice.

    1. mkl · · focus · HN ↗
      There's some info in the README, including:

      > Coder: a coding version with half of the experts removed. It reaches 91% of the full model's SWE-bench Verified score (measured by its authors) and fits 32 GB of RAM.

      <a href="https:&#x2F;&#x2F;github.com&#x2F;Niko1221&#x2F;Strata#which-model-should-i-pick" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;Niko1221&#x2F;Strata#which-model-should-i-pick

      1. nicce · · focus · HN ↗
        I wonder how this Coder compares to Qwen 3.8 27B. Can it be really better since they are competitive for same memory requirements?
        1. kennywinker · · focus · HN ↗
          125b at q2 is ~80gb

          27b at q4 is ~16gb

          So from a raw amount of data, qwen3.8-flash-next wins easily. But flash-next is an MoE model, so it only has 6b parameters active per token, vs 27b&#x27;s dense 27b per token. So 27b@q4 uses ~16gb of weights per token, and flash-next uses about 4gb of weights (125&#x2F;80 * 6).

          But those numbers don&#x27;t really tell us anything useful, because there is an interplay between total model size and active parameters and intelligence that isn&#x27;t obvious or simple.

          (sizes are based on the unsloth quants, not the coder variant, but the idea holds - this isn&#x27;t calculatable with simple math, you gotta test them and see)

        2. XCSme · · focus · HN ↗
          In my tests they do quite similarly, but 3.8 flash next is considerably (2x) more efficient and faster to respond.

          Both 27b and flash next are more stable on &quot;low&quot; reasoning, only for generative &#x2F;creative tasks, xhigh could be better, but both suffer from way too much reasoning at xhigh. And neither really support high, so low is the best reasoning effort.

          [0]: <a href="https:&#x2F;&#x2F;aibenchy.com&#x2F;compare&#x2F;qwen-qwen3-8-27b-low&#x2F;qwen-qwen3-8-flash-next-low&#x2F;" rel="nofollow">https:&#x2F;&#x2F;aibenchy.com&#x2F;compare&#x2F;qwen-qwen3-8-27b-low&#x2F;qwen-qwen3...

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.