‹ BackHN Continuity

Thread

Kolibri: A Sovereign Open-Weight Model

673 points · 327 comments · bastitx

  1. petesergeant · · focus · HN ↗
    I wish nothing but luck for an EU model, but:

    > intellectual-property safety

    My suspicion is that you simply can't build an even slightly competitive model without liberally stealing your training data, in 2026, as much as I'd like it to be otherwise. You can get to the point that I suspect most of the frontier labs are at, where you've laundered the initially stolen data through the creation of huge amounts of derivative synthetic data, but still. Anyone who isn't comfortable stealing their training data is bringing a knife to a gun fight, and is going to die a noble but inevitable death.

    1. befelix · · focus · HN ↗
      Disclaimer: I am part of the team that trained Kolibri, opinions are mine.

      You're right that especially big models benefit from training on copyrighted material in terms of world knowledge (especially from books). However, in the small model space imho agentic capabilities where the model looks up knowledge on the fly are much more important. That's what we focused on quite a bit during training. Personally, I also don't think stealing stuff is okay.

      1. petesergeant · · focus · HN ↗
        I hope you're right, but I guess we'll see when we get independent benchmarks. My intuition is that even for small models and models mostly focused on tool calling, you _still_ need all that extra contextual stuff for the magic, but I am further from the coalface than you are.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.