‹ BackHN Continuity

Thread

Kolibri: A Sovereign Open-Weight Model

671 points · 327 comments · bastitx

  1. miellaby · · focus · HN ↗
    The paper explains absolutely everything as if it was a tutorial &quot;how to made your own modern agentic LLM&quot;. They even tell how they made their dataset. <a href="https:&#x2F;&#x2F;aleph-alpha.com&#x2F;downloads&#x2F;tech-report.pdf" rel="nofollow">https:&#x2F;&#x2F;aleph-alpha.com&#x2F;downloads&#x2F;tech-report.pdf ; It&#x27;s the first time I see this level of openness.
    1. ivo-42 · · focus · HN ↗
      I worked on Kolibri, in particular pre-training data and mid-training. We strive to be as open as possible. Glad you like it.
      1. brcmthrowaway · · focus · HN ↗
        How do you cleanse the data at this scale?
        1. ivo-42 · · focus · HN ↗
          By various forms of deduplication (exact, fuzzy, substring), heuristic filters and distilling quality classifiers that annotate our data. Synthetic rephrases can also be considered a form of cleaning&#x2F;getting more out of existing noisy data.

          We have a lot of details in the tech report if you want to go deeper.

          1. stephantul · · focus · HN ↗
            Hey! I’m curious if you tried comparing luxical to model2vec classifiers for the pretraining.

            I’m one of the authors of model2vec, and working on training classifiers for this. I think model2vec could be better, but I haven’t had the opportunity to try this at scale. So if you did, knowing about it would be helpful!

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.