‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

741 points · 330 comments · snehesht

  1. hecturchi · · focus · HN ↗
    - Tiny context size or hours to load it

    - Hard to benefit from thinking and preserve thinking given token cost.

    - Low quants reduce accuracy heavMTP draft can make it make the same mistakes all the time when calling tools, formatting output or following basic guidelines. Otherwise 2x-4x slower.

    - K/V quants probably quantized too make things less accurate.

    Useful would be combinations with:

    - Full context size so it can code and think a bit.

    - Draft MTP <= 2 so it doesn't trip

    - Q4 quants or better so its accurate

    - q8 cache or better so it stays accurate.

    - 20 token/s so it finishes while reviewing previous step.

    - 1000 tokens/s context load so compactions don't waste 10+ minutes.

    - And enough left RAM for 50+ context checkpoints so that it can progress quuckly.

    Closest you have is Qwen3.6-35B-A3B-MTP.

    Latest gens (Qwen3.8 and co.) are just too big for low specs. 27B dense models seem to be ok for integrated >=92 GiB RAM.

    Source: I have low specs and tried them all for agentic use + coding.

    1. ohyes · · focus · HN ↗
      I’m using qwen3.8 quants (q3) effectively on a 5070 ti. A lot of it is about guardrails, but you also have to figure out how far (and in what ways) you can push a given model.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.