‹ BackHN Continuity

Thread

Kolibri: A Sovereign Open-Weight Model

671 points · 327 comments · bastitx

  1. petesergeant · · focus · HN ↗
    I wish nothing but luck for an EU model, but:

    > intellectual-property safety

    My suspicion is that you simply can't build an even slightly competitive model without liberally stealing your training data, in 2026, as much as I'd like it to be otherwise. You can get to the point that I suspect most of the frontier labs are at, where you've laundered the initially stolen data through the creation of huge amounts of derivative synthetic data, but still. Anyone who isn't comfortable stealing their training data is bringing a knife to a gun fight, and is going to die a noble but inevitable death.

    1. Zambyte · · focus · HN ↗
      Is it noble? The entire notion that training data can be "stolen" at all is quite silly. If I "steal" content that someone created to use for training, what am I actually stealing? They didn't lose anything. They still have everything they had before. What was "stolen" was "unrealized profit", or put another way: money that wasn't theirs, that they had no entitlement to. The only actual crime that is committed is "unauthorized copying", not stealing. Support and enforcement of copyright feels wildly authoritarian. It's hard to see it as noble.
      1. pepperoni_pizza · · focus · HN ↗
        That's fair, but then people like you complain when someone "steals" I mean distills openai or anthropic models.
        1. Zambyte · · focus · HN ↗
          I don't complain about that. Model distillation is excellent.
        2. brookst · · focus · HN ↗
          It’s poor form to argue against someone by imagining something totally different that they might believe, which would then make them hypocritical.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.