‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

806 points · 356 comments · snehesht

  1. hecturchi · · focus · HN ↗
    - Tiny context size or hours to load it

    - Hard to benefit from thinking and preserve thinking given token cost.

    - Low quants reduce accuracy heavMTP draft can make it make the same mistakes all the time when calling tools, formatting output or following basic guidelines. Otherwise 2x-4x slower.

    - K/V quants probably quantized too make things less accurate.

    Useful would be combinations with:

    - Full context size so it can code and think a bit.

    - Draft MTP <= 2 so it doesn't trip

    - Q4 quants or better so its accurate

    - q8 cache or better so it stays accurate.

    - 20 token/s so it finishes while reviewing previous step.

    - 1000 tokens/s context load so compactions don't waste 10+ minutes.

    - And enough left RAM for 50+ context checkpoints so that it can progress quuckly.

    Closest you have is Qwen3.6-35B-A3B-MTP.

    Latest gens (Qwen3.8 and co.) are just too big for low specs. 27B dense models seem to be ok for integrated >=92 GiB RAM.

    Source: I have low specs and tried them all for agentic use + coding.

    1. qeternity · · focus · HN ↗
      > Flash attention with who knows how much draft ensures it makes the same mistakes all the time and cannot call tools, format output or follow basic guidelines reliably.

      > Draft MTP <= 2 so it doesn't trip

      I am not sure you understand what either of these things do.

      Do you think that FA or MTP are lossy?

      1. hecturchi · · focus · HN ↗
        I mixed FA and MTP wrong in my original post, thanks for pointing it out.

        My experience is that draft-mtp=2 gives 25% improvement in tokens/s but the model is unable to call tools with the right arguments reliably. I have since then gone for smaller quants at draft-mtp=1 and that problem is at least gone so far.

        An agent needs to repeat tool calls in the right way, so cannot penalize repetition. This however leads to the agent retrying the wrong calls constantly.

        1. kadoban · · focus · HN ↗
          Something sounds wrong there. MTP should be impossible for it to degrade quality, it's exactly the same token stream.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.