‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

741 points · 330 comments · snehesht

  1. jacquesm · · focus · HN ↗
    LLM threads the world over are spammed with Strata links, it remains to be seen how much of the breathless hype remains standing once the honeymoon period is over. I've tried it but so far I have not seen anything that overly impressed me in terms of accuracy, though the speed is definitely there. I'm sure there are applications for LLMs where the quality of the answers is less important but I don't have any of those. YMMV.
    1. brcmthrowaway · · focus · HN ↗
      What does this inference engine do that others don't?
      1. jacquesm · · focus · HN ↗
        It is faster than comparable engines using the same model, but they use a lot of short-cuts.
    2. kristopolous · · focus · HN ↗
      I'm sick of these 4 day old vibe coded projects becoming hyped to hell while I spend a year on something and it goes nowhere

      What on God's green earth are these children doing?

      Do I need to get on tiktok? Do a dance? Make memes?

      I don't know who this kid is but I've looked at the code, it's all Claude. Is it discord? LinkedIn?

      The thing doesn't actually work very well. It's not even good software.

      So solid engineering, good documentation, functional software, that's all irrelevant noise.. There's some other magical handwaving Internet meme bullshit I'm just not getting

      I want to do high quality work but apparently I should be dicking around on social media

      1. jacquesm · · focus · HN ↗
        There are quite a few of these, indeed. Most of them either start of with llama.cpp (everybody's favorite to rip off, make a minor improvement to and try to establish a reputation) or vllm. But some of these actually do have something to bring to the table, for instance, by limiting the themselves to be able to focus on a smaller set of possibilities and this can lead to favorable outcomes for a smaller size of the audience. One particularly good example of this I think is 'ninfer' which mainly focuses on the Qwen family of models served up on RTX 5090's, which it does outstandingly well. Other people then go and fork that to adapt to their hardware so now there are 3090 and 4090 forks of ninfer.

        Then there is the 'unified memory' branch of inference engines, the most notable of which is probably DwarfStar 4 by 'Antirez', which also started off with a lot of code from the llama.cpp codebase.

        And then there is 'the rest', but even there, some of these have interesting bits and I always hope that eventually those bits will make their way back to the engine where it started.

        I run both ninfer and llama.cpp, vllm is an unmaintainable mess even though there usually is a performance edge (it is great if you are serving up for commercial purposes so you can tweak it for one set of hardware and one particular model). You'd essentially need to dedicate a week or more to getting a new model up and running on a particular set of hardware if it does not nicely match with the recipes found online.

        One exception is the DGX Spark series, there vllm is supported by the manufacturer and the hardware is very consistent from one box to another. But the performance isn't really there when compared to a fat PC with a bunch of GPUs. (The 200 GB/s memory bandwidth is a serious performance killer).

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.