‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

741 points · 330 comments · snehesht

  1. jacquesm · · focus · HN ↗
    LLM threads the world over are spammed with Strata links, it remains to be seen how much of the breathless hype remains standing once the honeymoon period is over. I've tried it but so far I have not seen anything that overly impressed me in terms of accuracy, though the speed is definitely there. I'm sure there are applications for LLMs where the quality of the answers is less important but I don't have any of those. YMMV.
    1. brcmthrowaway · · focus · HN ↗
      What does this inference engine do that others don't?
      1. jacquesm · · focus · HN ↗
        It is faster than comparable engines using the same model, but they use a lot of short-cuts.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.