‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

741 points · 330 comments · snehesht

  1. mark_l_watson · · focus · HN ↗
    Qwen 3.8 Flash Next is amazing. I only have a 64G Mac so I have to run Sushi project’s 3 bit quant. Amazing results with pi-dev. More for fun than anything else, but I am trying to do as much as possible with local models, now rarely falling back to a paid deepseek-4.1-flash API.

    Progress on running local models has been amazing.

    1. fsiefken · · focus · HN ↗
      Yes, I am running the same on a 64G mc. It's good, but slow at 25 tps on average! I want > 100 tps - but I don't have $5k to spare for an m5 ultra or an nvidia setup.

      So I still wonder if one could get good enough quality with a faster higher quant or superoptimized Qwen3.8-27b with dflash2

      <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;nathansutton&#x2F;Qwen3.8-27B-Ternary-Bonsai-2-DFlash2-MLX" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;nathansutton&#x2F;Qwen3.8-27B-Ternary-Bons...

      or a MoE retrofit like Qwen3.8-35B-A3B with or without mtp

      <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;NovaeonStudio&#x2F;Qwen3.8-35B-A3B-Distill-Heretic-oQ8-fp16-mtp" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;NovaeonStudio&#x2F;Qwen3.8-35B-A3B-Distill...

      <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;IsValorum&#x2F;Qwen3.8-35B-A3B-Distill-MLX-APEX-MiniPlus-Abliterated" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;IsValorum&#x2F;Qwen3.8-35B-A3B-Distill-MLX...

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.