Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
Unofficial Hacker News client; not affiliated with Y Combinator.
a11r · · focus · HN ↗
Winfred-zz · · focus · HN ↗
│------------------- │ Ninfer-3090 │ Strata
│ Code generation │ 52/78 (66.7%) │ 70/78 (89.7%)
│ Code completion │ 40/50 (80.0%) │ 44/50 (88.0%)
│ Total------------- │ 92/128 (71.9%) │ 114/128 (89.1%)
│ API failures------ │ 10 │ 5
- Ninfer generation: ~122 min total.
- Strata generation: ~142 min total.
So strata is a little slower, but keep in mind that ninfer-3090 is very optimized for a Qwen 3.8. Standard Qwen 3.8 runs at 20 t/s, this modified version can do 50 t/s (but it's extremely long in it's thinking, it just goes on and on.
This is on a 3090 that will crash unless power capped, with a Zen 2 CPU, 64GB DDR4 with a PCIe that refuses to go higher than 8x (basically pretty crappy all in all).
Yet with some tweaking and optimizing I still manage to get strata to run at 40 to 60 t/s.
That strata has been optimized on my Oh My Pi conversations. So when I'm using it, it's probably faster and closer to ninfer in speed than during those unoptimized benchmark tests.
seviu · · focus · HN ↗
For me, I didnt know I was on 4x and I was able to, giving it enough permissions, with a good harness like PI, to figure things out and go 4x -> 8x -> 16x.
It probably will take some opening the case and switching ssds around, but it's amazing how powerful these models have become.