‹ BackHN Continuity

Thread

Kolibri: A Sovereign Open-Weight Model

671 points · 327 comments · bastitx

  1. petesergeant · · focus · HN ↗
    I wish nothing but luck for an EU model, but:

    > intellectual-property safety

    My suspicion is that you simply can't build an even slightly competitive model without liberally stealing your training data, in 2026, as much as I'd like it to be otherwise. You can get to the point that I suspect most of the frontier labs are at, where you've laundered the initially stolen data through the creation of huge amounts of derivative synthetic data, but still. Anyone who isn't comfortable stealing their training data is bringing a knife to a gun fight, and is going to die a noble but inevitable death.

    1. rpdillon · · focus · HN ↗
      This doesn't seem to be true. There's a clear legal path via the first-sale doctrine to train models on copyrighted works. It's been years now, and publishers still don't seem to be offering anything for training (e.g. bulk licenses solely for training use), but adversarial interoperability via cutting up books and scanning them remains perfectly legal.

      There's also the ability to distill other models, which is also not illegal (though I'm sure they like to come after whomever for TOS violations, but thats a civil matter).

      And, of course, the obligatory copying-isn't-theft observation. A recent supreme court judgment put it well.

      > Since the statutorily defined property rights of a copyright holder have a character distinct from the possessory interest of the owner of simple “goods, wares, [or] merchandise,” interference with copyright does not easily equate with theft, conversion, or fraud. The infringer of a copyright does not assume physical control over the copyright, nor wholly deprive its owner of its use. Infringement implicates a more complex set of property interests than does run-of-the-mill theft, conversion, or fraud.

      Folks are pretty smart here, I think we can handle these nuances, even if we don't agree about whether they are good.

      Edit: reading through the full text of their post, it looks like they are using common crawl, which is likely just as much of a copyright infringement as Anna's Archive -- it's not like published works have a unique claim to copyright. I think this strengthens your point, though: I was expecting to see scans as training data, but it doesn't appear to be the case.

      1. mapontosevenths · · focus · HN ↗
        "The Congress shall have Power To ... promote the Progress of Science and useful Arts, by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries." - The United States Constitution

        Copyright is a government mandated monopoly that was only granted in order to advance the arts and science. Any interpretation that runs contrary to that is bollocks being used by the religiously or financially motivated to serve their own petty interests to the detriment of societies.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.