Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
Unofficial Hacker News client; not affiliated with Y Combinator.
sleight42 · · focus · HN ↗
Maybe there's some whackadoodle way to only use a portion of the VRAM for experts for one request and another portion for experts for the other such that requests could actually run parallel instead of concurrently?
EDIT: I'm wrong! It already can do this with the parallel argument.
However, a few days ago, the dev(s now?) added swapping contexts to and from RAM.
Also, from my above, I don't see wh