‹ BackHN Continuity

Thread

Kolibri: A Sovereign Open-Weight Model

668 points · 327 comments · bastitx

  1. peterBlue75 · · focus · HN ↗
    The thing to note here, besides the transparency and the fact that it’s actually a good model that also works well on coding and agentic tasks, is that it’s the first release by a team formed less than a year ago, with a strong focus on iteration velocity. There’s more to come.

    disclaimer: I‘m part of the training team, happy to answer any questions

    1. ducktective · · focus · HN ↗
      - Is it possible to train only on math and logic materials and expect the model's response in math questions to be superior to general models with the same training/inference compute hardware?

      - Are there non-LLM approaches to the above task with the goal of achieving a non-hallucinatory agent?

    2. sspiff · · focus · HN ↗
      Any future plans you can share to further evolve / train models in this size class?

      It seems like a very suitable size for local AI models on reasonably high end consumer devices, given it's low active parameter count and a mixed 8bit/4bit quant would fit easily inside 64GB of memory.

      1. peterBlue75 · · focus · HN ↗
        I can't comment on concrete sizes of future models, partly because it's not decided yet. But: Pre-training for Kolibri only started in August so there's a high chance continued post-training - which we plan to do - will yield some nice checkpoints. We believe this size and sparsity allows achieving great inference throughput at reasonable levels of intelligence.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.