‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

806 points · 356 comments · snehesht

  1. hecturchi · · focus · HN ↗
    - Tiny context size or hours to load it

    - Hard to benefit from thinking and preserve thinking given token cost.

    - Low quants reduce accuracy heavMTP draft can make it make the same mistakes all the time when calling tools, formatting output or following basic guidelines. Otherwise 2x-4x slower.

    - K/V quants probably quantized too make things less accurate.

    Useful would be combinations with:

    - Full context size so it can code and think a bit.

    - Draft MTP <= 2 so it doesn't trip

    - Q4 quants or better so its accurate

    - q8 cache or better so it stays accurate.

    - 20 token/s so it finishes while reviewing previous step.

    - 1000 tokens/s context load so compactions don't waste 10+ minutes.

    - And enough left RAM for 50+ context checkpoints so that it can progress quuckly.

    Closest you have is Qwen3.6-35B-A3B-MTP.

    Latest gens (Qwen3.8 and co.) are just too big for low specs. 27B dense models seem to be ok for integrated >=92 GiB RAM.

    Source: I have low specs and tried them all for agentic use + coding.

    1. ranger_danger · · focus · HN ↗
      What about Bonsai 2? You can fit Qwen3.8 27B on an 8GB GPU with it, and upstream llama.cpp support is already being worked on (they just got System1 support too).
      1. arcanemachiner · · focus · HN ↗
        The ternary model? Hopefully those are worth a damn in a few years, but currently just an interesting toy from what I understand.
      2. Aurornis · · focus · HN ↗
        The Bonsai models are really bad when you actually use them for more than short responses.

        Their marketing made it look like a breakthrough, but in my experience it’s just the next step down from the Q2 quants in both size and quality.

        Q2 quants are already not very useful in my experience. The Bonsai models are even worse.

        If you only need 80% plausible outputs that don’t need to reference a lot of context they can be useful. If you try to use them for real tasks it feels like time warping back to 2023 when you LLMs were barely useful if you babysat every word of the output.

        1. ranger_danger · · focus · HN ↗
          Have you actually used Bonsai 2 though and not just the original Bonsai? The experience is vastly improved but still requires a custom llama.cpp fork to use as of right now.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.