‹ BackHN Continuity

Thread

Kolibri: A Sovereign Open-Weight Model

668 points · 327 comments · bastitx

  1. miellaby · · focus · HN ↗
    The paper explains absolutely everything as if it was a tutorial &quot;how to made your own modern agentic LLM&quot;. They even tell how they made their dataset. <a href="https:&#x2F;&#x2F;aleph-alpha.com&#x2F;downloads&#x2F;tech-report.pdf" rel="nofollow">https:&#x2F;&#x2F;aleph-alpha.com&#x2F;downloads&#x2F;tech-report.pdf ; It&#x27;s the first time I see this level of openness.
    1. zelphirkalt · · focus · HN ↗
      In my opinion not being open about which data is ingested and trained on, and trying to make that a repeatable thing for a third party, is not worth being called &quot;open&quot;. Glad they did that.
      1. slow_typist · · focus · HN ↗
        How, is the training data public?
      2. davidjfelix · · focus · HN ↗
        to be fair, this discussion has been had numerous times here and the industry has arrived on &quot;open-weight&quot; to describe the practice of releasing the post-training weights in an open manner but not releasing the data it was trained on.

        That&#x27;s what they call this and I think it&#x27;s a pretty clear definition these days to people in the industry.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.