‹ BackHN Continuity

Thread

Pirate Face Rescues LLM Models from Deletion

569 points · 149 comments · skepticalgenius

  1. phoyd · · focus · HN ↗
    Torrents should really be the preferred method for distributing AI model weights. Why rely on a single point of failure like Hugging Face? BitTorrent was made for exactly this.
    1. CodesInChaos · · focus · HN ↗
      In my experience public torrents often die as they grow older. It doesn't help that BitTorrent V1 makes long term seeding annoying, and BitTorrent V2 is almost never used.
      1. monsieurbanana · · focus · HN ↗
        I never understood this, is there anything that makes it difficult for the original uploader, the one that supposedly offers the file directly, to offer a torrent instead for the same amount of time?

        As far as perennity is concerned it seems strictly better.

        1. zenoprax · · focus · HN ↗
          Every change to the source is effectively a new torrent. This creates a ton of fragmentation as data is reorganized, remixed, reencoded, and so on.

          You can see this with many Linux distros: there is no single Debian torrent that people seed for years because there's always a refreshed version.

          Distros are a bad use case for P2P anyway since you depend on upstream as soon as you start upgrading and installing packages.

          1. vova_hn2 · · focus · HN ↗
            IPFS has a solution [0] to this problem

            [0] <a href="https:&#x2F;&#x2F;specs.ipfs.tech&#x2F;ipns&#x2F;ipns-record&#x2F;" rel="nofollow">https:&#x2F;&#x2F;specs.ipfs.tech&#x2F;ipns&#x2F;ipns-record&#x2F;

            1. mitxela · · focus · HN ↗
              Most of IPFS doesn&#x27;t actually work very well, if you&#x27;ve ever tried to use it
              1. vova_hn2 · · focus · HN ↗
                &gt; if you&#x27;ve ever tried to use it

                Heh, you got me :) IPFS is one of those things that I love reading and about and thinking about using someday, but somehow never get around to it.

                1. mitxela · · focus · HN ↗
                  The problems begin with taking 5-10 minutes to locate a file on the network. That&#x27;s right, when you ask for a file it takes 5-10 minutes. Also if the file isn&#x27;t in the network at all then it never terminates.

                  Nobody noticed because everyone just used the central web gateway that cached every file anyone ever accessed.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.