Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
Unofficial Hacker News client; not affiliated with Y Combinator.
mark_l_watson · · focus · HN ↗
Progress on running local models has been amazing.
generalizations · · focus · HN ↗
fsiefken · · focus · HN ↗
So I still wonder if one could get good enough quality with a faster higher quant or superoptimized Qwen3.8-27b with dflash2
<a href="https://huggingface.co/nathansutton/Qwen3.8-27B-Ternary-Bonsai-2-DFlash2-MLX" rel="nofollow">https://huggingface.co/nathansutton/Qwen3.8-27B-Ternary-Bons...
or a MoE retrofit like Qwen3.8-35B-A3B with or without mtp
<a href="https://huggingface.co/NovaeonStudio/Qwen3.8-35B-A3B-Distill-Heretic-oQ8-fp16-mtp" rel="nofollow">https://huggingface.co/NovaeonStudio/Qwen3.8-35B-A3B-Distill...
<a href="https://huggingface.co/IsValorum/Qwen3.8-35B-A3B-Distill-MLX-APEX-MiniPlus-Abliterated" rel="nofollow">https://huggingface.co/IsValorum/Qwen3.8-35B-A3B-Distill-MLX...