The absence of any comparison to Qwen3.8 Flash, another MoE model with a small-ish (6B) number of active parameters, is pretty striking. Instead, it's compared with Qwen3-Next 80B-A3B, a model released almost a full year ago.
I get that doesn't invalidate the real "point" of the model, but...
Can't you infer the comparisons you would like from baseline results provided? There are better results in the sibling post on HN: <a href="https://aleph-alpha.com/en/blog/kolibri-has-landed-a-sovereign-open-weight-model/" rel="nofollow">https://aleph-alpha.com/en/blog/kolibri-has-landed-a-soverei...
Qwen3.8 27B scored notably higher in most of the provided benchmarks, including the German-specific ones. The only "downside" is that inference is much more costly and slow, since it's a dense model.
Qwen3.8 Flash-Next appears to usually "benchmark higher" than 27B, while remaining fast.
I'm sure I could dig up the equivalent benchmarks for Flash and do the comparison myself, but as far as inference goes, it's messy. Consider that Qwen3.5 35B-A3B scores higher than Qwen3.6 on some of the German-specific benchmarks.
So it seems superficially plausible that Qwen3.8 Flash-Next might not be "27B but faster" in the ways that are important for this model. Or it could just "be superior" in all ways.
Either way, I don't think an LLM has to be "the best" at anything to be worthwhile, necessarily. And I kind of distrust benchmarks on top of that, so...
spijdar · · focus · HN ↗
I get that doesn't invalidate the real "point" of the model, but...
okamiueru · · focus · HN ↗
spijdar · · focus · HN ↗
Qwen3.8 27B scored notably higher in most of the provided benchmarks, including the German-specific ones. The only "downside" is that inference is much more costly and slow, since it's a dense model.
Qwen3.8 Flash-Next appears to usually "benchmark higher" than 27B, while remaining fast.
I'm sure I could dig up the equivalent benchmarks for Flash and do the comparison myself, but as far as inference goes, it's messy. Consider that Qwen3.5 35B-A3B scores higher than Qwen3.6 on some of the German-specific benchmarks.
So it seems superficially plausible that Qwen3.8 Flash-Next might not be "27B but faster" in the ways that are important for this model. Or it could just "be superior" in all ways.
Either way, I don't think an LLM has to be "the best" at anything to be worthwhile, necessarily. And I kind of distrust benchmarks on top of that, so...