Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
Unofficial Hacker News client; not affiliated with Y Combinator.
snehesht · · focus · HN ↗
<a href="https://huggingface.co/Qwen/Qwen3.8-Flash-Next" rel="nofollow">https://huggingface.co/Qwen/Qwen3.8-Flash-Next
roscas · · focus · HN ↗
This is not a very fast desktop. Memory speed is around 2000mhz only. My SSD is some of the worst SSD I've seen and 3080 had its days of glory.
I still have code, chromium, librewolf and many other programs running. I have video streams running while I also watch tv and many times youtube videos.
I use it with the browser that has a great dashboard and with hermes agent and that it really makes this amazing.Only change I made is to set thinking to low.
This is a coding model. Any other task, I still use Ornith 1.5 35B that throws 20t/sec and Laguna.XS-2.0.
StumpChunkman · · focus · HN ↗
Abishek_Muthian · · focus · HN ↗
I run small models on various kind of devices including 1.7B model on the original Jetson Nano (4B) abandoned by Nvidia, I had to upscale the software (OS, Lllama.cpp etc.) to run the model but the model runs at 17t/s.
Small models are great at NLP stuff like classification (e.g. bookmarking), ASR etc. I have built custom browser extensions to save time with bookmarking and categorizing to use with local LLMs.