Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
Unofficial Hacker News client; not affiliated with Y Combinator.
jacquesm · · focus · HN ↗
kristopolous · · focus · HN ↗
I don't know who this kid is but I've looked at the code, it's all Claude. Saying you can run on 2-bit quant with some optimization flags at 100t/s isn't like some "oh my goodness" ... it's wasting my time with noise.
What are they doing? Is it Tiktok? Discord? LinkedIn?
I want to do high quality work but apparently I should be dicking around on social media
jacquesm · · focus · HN ↗
Then there is the 'unified memory' branch of inference engines, the most notable of which is probably DwarfStar 4 by 'Antirez', which also started off with a lot of code from the llama.cpp codebase.
And then there is 'the rest', but even there, some of these have interesting bits and I always hope that eventually those bits will make their way back to the engine where it started.
I run both ninfer and llama.cpp, vllm is an unmaintainable mess even though there usually is a performance edge (it is great if you are serving up for commercial purposes so you can tweak it for one set of hardware and one particular model). You'd essentially need to dedicate a week or more to getting a new model up and running on a particular set of hardware if it does not nicely match with the recipes found online.
One exception is the DGX Spark series, there vllm is supported by the manufacturer and the hardware is very consistent from one box to another. But the performance isn't really there when compared to a fat PC with a bunch of GPUs. (The 200 GB/s memory bandwidth is a serious performance killer).