‹ BackHN Continuity

Thread

Japanese used bookstores see 5x sales surge as books are being bought by the ton

96 points · 130 comments · speckx

  1. zirkonit · · focus · HN ↗
    Our family volunteers at a nonprofit that moves a huge number of books. We take donations and run massive charity sales, clearing tens of thousands of books a month. Pricing works like a ladder: you try to sell a book for a couple of bucks, then for a dollar, then by the $5 bag, then for free, and you still end up with thousands of books nobody wants even at no cost. These used to go straight to pulp. Now they go to AI labs for scanning. Would we rather they were read, or at least owned, by someone? Yes. Is scanning better than turning them into toilet paper? Yes, even if only marginally.

    I'm sure the Anthropics of the world have shredded the last known copies of plenty of titles. But nobody talks about the last known copies that get pulped every day simply because there are so many books nobody wants and storing them is expensive.

    1. sly010 · · focus · HN ↗
      Well, they could turn the bad faith story into a good faith story by making them available for everyone to download perhaps. (AI companies "saving" old books!) But that would require giving a s*t which they don't and that is the real problem imho.
      1. Aurornis · · focus · HN ↗
        They legally cannot do this.
        1. theroadnotbacon · · focus · HN ↗
          That certainly hasn’t stopped them before… IP theft is kind of their whole thing, isn’t it?
          1. Aurornis · · focus · HN ↗
            Distributing copyrighted works (prior to expiration of their copyright) verbatim is illegal.

            Training an LLM on copyrighted works is not illegal.

            This whole debate has been tried in court already. Calling it IP theft only stands on individual moral grounds, but the law allows for derivative works.

          2. syrrim · · focus · HN ↗
            Google attempted to do this 15 years ago, they got sued and stopped. It turns out that tech companies occasionally do have to follow the law, you'd think people would be happier about that...
            1. keeda · · focus · HN ↗
              Wait, if you’re talking about the Google Books case, Google won. Maybe they made adjustments on how they served results but they certainly did not stop.
              1. ndiddy · · focus · HN ↗
                The Google Books settlement was originally going to make Google into a clearinghouse for scans of out-of-print books. The scans would have been available for individuals to purchase for a reasonable price, and libraries and institutions would have been able to subscribe to a service that would give patrons access to the full text of every book. This deal fell apart because some research libraries and authors argued this was anti-competitive, as anyone wanting to make a competing service would have to go through the same process as Google of settling a class action lawsuit. They instead wanted Congress to pass a law to free up the rights to orphaned books. No progress on this was subsequently made because nobody in Congress cares enough about the rights to out-of-print books to get legislation passed. The whole reason why they're out of print when ebooks and print-on-demand exist is that they won't get enough sales to make it worth the time and money to figure out who the royalties should go to. The result is that nobody outside Google gets to see the full Google Books scans.
                1. shagie · · focus · HN ↗
                  Those scans are held at HathiTrust Research Center <a href="https:&#x2F;&#x2F;www.hathitrust.org&#x2F;about&#x2F;research-center&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.hathitrust.org&#x2F;about&#x2F;research-center&#x2F; <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;HathiTrust" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;HathiTrust

                  &gt; HathiTrust Digital Library is a large-scale collaborative repository of digital content from research libraries, administered by the University of Michigan. Its holdings include content digitized via Google Books and the Internet Archive digitization initiatives, as well as content digitized locally by libraries.

                  <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Authors_Guild,_Inc._v._HathiTrust" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Authors_Guild,_Inc._v._HathiTr... and <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Authors_Guild,_Inc._v._Google,_Inc.#Impact" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Authors_Guild,_Inc._v._Google,...

                  &gt; Authors Guild, Inc. v. HathiTrust (2014) was a following case related to HathiTrust, a project by the libraries of the Big Ten Academic Alliance and the University of California systems that combined their digital library collections with those of Google&#x27;s Book Search. The HathiTrust case differed in two primary factors which were raised by the plaintiffs: that for viewers with disabilities, they could view the scanned text through a screen reader to make it easier to read, and offering to print out the scans as replacement copies for members of the universities if they could verify their original copies were lost or damaged. Both uses were deemed also to be fair use by the Second Circuit.

                  &gt; The subject of the copyright of orphan works – works that may still be under copyright but with no identifiable rights holder – was a significant point of debate after both this and HathiTrust. Normally, libraries have been hesitant to loan digital copies of orphaned works as libraries may be liable for copyright violations should the copyright owner step forward to claim ownership.

                  The bill on orphan works that didn&#x27;t pass was <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Shawn_Bentley_Orphan_Works_Act_of_2008" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Shawn_Bentley_Orphan_Works_Act...

                  <a href="https:&#x2F;&#x2F;www.hathitrust.org&#x2F;the-collection&#x2F;search-access&#x2F;copyright-access&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.hathitrust.org&#x2F;the-collection&#x2F;search-access&#x2F;copy...

                  And there are exceptions for copyrighted works allowing them to lend them out.

                  &gt; Protected by copyright law, but made available: Protected by copyright law but made available on a strictly limited basis in accordance with the statutory limitations including, but not limited to, Section 107 provisions for fair use, Section 108 provisions for libraries and archives, and the rights provided to registered users with disabilities. In the absence of an applicable exception, no further reproduction or distribution is permitted by any means without the permission of the copyright holder. Lawful uses of works are provided only under the following conditions ...

                  Scanning still continues. <a href="https:&#x2F;&#x2F;www.hathitrust.org&#x2F;member-libraries&#x2F;contribute-content&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.hathitrust.org&#x2F;member-libraries&#x2F;contribute-conte... - though it&#x27;s not at the same rate as it was during google books project.

        2. sly010 · · focus · HN ↗
          That would require effort (to sort, acquire copyright, etc) which they wouldn&#x27;t put in. Because they don&#x27;t care.

          People obviously feel bad about companies doing this. People reading these stories don&#x27;t care what&#x27;s legal, they care what&#x27;s ethical. Heck, re-publishing long lost material would make AI companies heroes instead of bad guys.

          1. Aurornis · · focus · HN ↗
            &gt; That would require effort (to sort, acquire copyright, etc) which they wouldn&#x27;t put in. Because they don&#x27;t care.

            I don&#x27;t think you have any idea how expensive it is to acquire the copyright for a single book with the intent of making it freely available online. That&#x27;s equivalent to asking the rights holders to perpetually forgo all possible earnings from the material, and they expect to be compensated accordingly. Even paying lawyers to begin assembling what&#x27;s needed to make this happen would be five figures per book to get started.

        3. 2OEH8eoCRo0 · · focus · HN ↗
          They legally cannot scan them in entirety either but they are.
          1. ChickeNES · · focus · HN ↗
            Again, under Bartz v Anthropic they can scan and train on whatever they want, as long as the original is lost in the process.
            1. 2OEH8eoCRo0 · · focus · HN ↗
              TIL thanks.
        4. wareya · · focus · HN ↗
          What I would do is announce that the scans are being preserved and will donated to the library of congress or whatever other institution is legally able to hold onto stuff like this. But that would more directly tie specific companies to the practice of destroying unwanted old books, so nobody&#x27;s going to do it, or even announce it.
          1. shagie · · focus · HN ↗
            The Library of Congress already has a copy of the book.
            1. wareya · · focus · HN ↗
              The library of congress does not have copies of random japanese books.
              1. shagie · · focus · HN ↗
                That would be National Diet Library <a href="https:&#x2F;&#x2F;www.ndl.go.jp&#x2F;en&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.ndl.go.jp&#x2F;en&#x2F;

                <a href="https:&#x2F;&#x2F;www.ndl.go.jp&#x2F;en&#x2F;collect&#x2F;deposit&#x2F;deposit" rel="nofollow">https:&#x2F;&#x2F;www.ndl.go.jp&#x2F;en&#x2F;collect&#x2F;deposit&#x2F;deposit

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.