Curious if or when we'll see the Qwen4 series, one thing I love with Qwen is it comes a much larger range of sizes so I can experiment which extremely small llms.
I don't think Qwen3.8-Omni-X will ever be released.
The last one was: Qwen3-Omni-30B-A3B <a href="https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct" rel="nofollow">https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct
And maybe Qwen4 won't be released, they only release Qwen3.8 27B (and a mostly unusable 125B). There are definitively slowing down open weight release.
I assume you are talking about qwen3.8-flash-next. Support for it on some places, like llama.cpp, is still wip (depending on configuration) but it looks like a very capable model in it's category.
Very capable yes but very slow.
27B is relatively easy to run, but the 125b one need around 128Gb of RAM (DDR4 isn't enough, you need DDR5 to be quick enough, that's $3000 alone, you also need a graphic card). DDR4 is caped @20tps.
So, the cost of a setup to run Qwen-Flash-Next at +40tks is around $3000. Too much for most people.
With only a RTX 4090, you will reach 30tps (with DDR5...), not +40tks, and it's about the limit to be usable. Oh ! I forget Apple device too, it's a good option to run this model I guess, but still slow.
Yet, as you said, it's still a wip implementation, it may improve soon (MTP support is about to be merged in llama.cpp soon).
tolugenius · · focus · HN ↗
_ache_ · · focus · HN ↗
The last one was: Qwen3-Omni-30B-A3B <a href="https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct" rel="nofollow">https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct
And maybe Qwen4 won't be released, they only release Qwen3.8 27B (and a mostly unusable 125B). There are definitively slowing down open weight release.
Iolaum · · focus · HN ↗
I assume you are talking about qwen3.8-flash-next. Support for it on some places, like llama.cpp, is still wip (depending on configuration) but it looks like a very capable model in it's category.
_ache_ · · focus · HN ↗
So, the cost of a setup to run Qwen-Flash-Next at +40tks is around $3000. Too much for most people.
With only a RTX 4090, you will reach 30tps (with DDR5...), not +40tks, and it's about the limit to be usable. Oh ! I forget Apple device too, it's a good option to run this model I guess, but still slow.
Yet, as you said, it's still a wip implementation, it may improve soon (MTP support is about to be merged in llama.cpp soon).