‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

741 points · 330 comments · snehesht

  1. esafak · · focus · HN ↗
    Has anyone calculated the effective intelligence of these quantized models?

    I think publishing benchmarks with quantized models should become standard practice.

    1. nsagent · · focus · HN ↗
      See this recent paper: Quantization Degradation in Large Language Models: A Signal–Noise Perspective [1].

        We observe that such degradation varies substantially across these factors: 4-bit quantization usually preserves performance, 2-bit often causes broad degradation
      
      This repo uses 2-bit quantization and removes some of the experts for its smallest fastest model. Make of that what you will.

      [1]: <a href="https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2608.08188" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2608.08188

      1. merbanan · · focus · HN ↗
        I created a pruned experts model of the q2 quant, while it gave good performance on limited hardware there was severe quality degradation.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.