‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

741 points · 330 comments · snehesht

  1. tracerbulletx · · focus · HN ↗
    The interesting thing here is that it's a model specialized fork of a generic inference engine that unlocks consumer hardware to run a bigger model with useable performance than it could before.
    1. conmod278 · · focus · HN ↗
      botspeak
      1. tracerbulletx · · focus · HN ↗
        The emergence of model specific inference (for consumers) getting big performance wins is way more worth while to talk about than random comments on what people think about the qwen family of models. Even the resource management of Strata is less interesting. I think it's likely we'll start seeing more hand/llm crafted inference for different architectures.
      2. robertkarl · · focus · HN ↗
        dang did his bot read this as an instruction to comment again?
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.