‹ BackHN Continuity

Thread

Kolibri: A Sovereign Open-Weight Model

668 points · 327 comments · bastitx

  1. spijdar · · focus · HN ↗
    The absence of any comparison to Qwen3.8 Flash, another MoE model with a small-ish (6B) number of active parameters, is pretty striking. Instead, it's compared with Qwen3-Next 80B-A3B, a model released almost a full year ago.

    I get that doesn't invalidate the real "point" of the model, but...

    1. okamiueru · · focus · HN ↗
      Can&#x27;t you infer the comparisons you would like from baseline results provided? There are better results in the sibling post on HN: <a href="https:&#x2F;&#x2F;aleph-alpha.com&#x2F;en&#x2F;blog&#x2F;kolibri-has-landed-a-sovereign-open-weight-model&#x2F;" rel="nofollow">https:&#x2F;&#x2F;aleph-alpha.com&#x2F;en&#x2F;blog&#x2F;kolibri-has-landed-a-soverei...
      1. spijdar · · focus · HN ↗
        Maybe?

        Qwen3.8 27B scored notably higher in most of the provided benchmarks, including the German-specific ones. The only &quot;downside&quot; is that inference is much more costly and slow, since it&#x27;s a dense model.

        Qwen3.8 Flash-Next appears to usually &quot;benchmark higher&quot; than 27B, while remaining fast.

        I&#x27;m sure I could dig up the equivalent benchmarks for Flash and do the comparison myself, but as far as inference goes, it&#x27;s messy. Consider that Qwen3.5 35B-A3B scores higher than Qwen3.6 on some of the German-specific benchmarks.

        So it seems superficially plausible that Qwen3.8 Flash-Next might not be &quot;27B but faster&quot; in the ways that are important for this model. Or it could just &quot;be superior&quot; in all ways.

        Either way, I don&#x27;t think an LLM has to be &quot;the best&quot; at anything to be worthwhile, necessarily. And I kind of distrust benchmarks on top of that, so...

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.