‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

741 points · 330 comments · snehesht

  1. Tepix · · focus · HN ↗
    All headlines about LLM performance MUST have the quantization also mentioned in the headline.

    You know, so you're not wasting your time like in this post.

    1. rpdillon · · focus · HN ↗
      Quants vary by model. DS4 is very credible at a 2-bit quant. Not sure about Qwen 3.8 Flash Next; I run it at a 4-bit quant and it's too slow, so I'm trying out DwarfStar today to see if that improves things.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.