‹ BackHN Continuity

Thread

Japanese used bookstores see 5x sales surge as books are being bought by the ton

93 points · 128 comments · speckx

Loading the complete thread in the background. This saved snapshot is available now. Refresh

  1. swingandamiss · · focus · HN ↗
    Time for the daily article on HN about Japan. The obsession with Japan on HN is very weird.
  2. duchanjo · · focus · HN ↗
    Could this be for AI training data?
    1. ErneX · · focus · HN ↗
      It’s on the headline if you visit the link.
  3. panny · · focus · HN ↗
    A dark age will come. AI shredders destroy all the books then hallucinate what they once contained.
    1. manarth · · focus · HN ↗
      Once upon a time, Hansel and Gretel were walking through the woods when they met a Sleeping Beauty called Snow White. As they tried to wake Beauty, a naked Emperor walked in screaming "Off with his head" before a Big Bad Wolf started huffing and puffing.
    2. aprilthird2021 · · focus · HN ↗
      I'm the opposite of an AI doomer but this is actually scary to me. Once knowledge is hollowed out like this how can we get it back?
      1. MarkusQ · · focus · HN ↗
        We could go out in the world, have experiences, cogitate upon them, learn to write well, and then do so? It ain't easy, but that's the way it used to be done.
      2. mistercheph · · focus · HN ↗
        This is unironically part of their business plan: don't just regurgitate the world's information but also destroy all other sources of information.
      3. shagie · · focus · HN ↗
        All the books that they are scanning exist in the national libraries where they were published and have been for the past century or so.

        Lets find a random conference book on Amazon... Genetics, Radiobiology and Radiology Proceedings, Mid-Western Conference. It was published in 1959. It's a rare book in that I can only find one copy of it on Amazon.

        <a href="https:&#x2F;&#x2F;www.amazon.com&#x2F;Radiobiology-Radiology-Proceedings-Mid-Western-Conference&#x2F;dp&#x2F;B000S5UTRG&#x2F;ref=sr_1_6" rel="nofollow">https:&#x2F;&#x2F;www.amazon.com&#x2F;Radiobiology-Radiology-Proceedings-Mi...

        Lets go find it...

        <a href="https:&#x2F;&#x2F;search.catalog.loc.gov&#x2F;instances&#x2F;e3fb7a94-2b25-5388-a8e8-41332625cf93?option=keyword&amp;query=Genetics%2C%20Radiobiology%20and%20Radiology%20Proceedings%2C%20Mid-Western%20Conference" rel="nofollow">https:&#x2F;&#x2F;search.catalog.loc.gov&#x2F;instances&#x2F;e3fb7a94-2b25-5388-...

        There&#x27;s one copy onsite at the Library of Congress in the General Collections and another copy offsite. You can get a reading card for the Science and Business Reading Room ( <a href="https:&#x2F;&#x2F;www.loc.gov&#x2F;research-centers&#x2F;science-and-business&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.loc.gov&#x2F;research-centers&#x2F;science-and-business&#x2F; ) and request that book and read it.

        It also happens that my alma mater has a copy of the book too. <a href="https:&#x2F;&#x2F;search.library.wisc.edu&#x2F;catalog&#x2F;999554514302121" rel="nofollow">https:&#x2F;&#x2F;search.library.wisc.edu&#x2F;catalog&#x2F;999554514302121 - it&#x27;s in the stacks in Ebling Library.

        Every book that has been published, there&#x27;s a copy of it somewhere. If it was published in the UK, it is in the British Library ( <a href="https:&#x2F;&#x2F;youtu.be&#x2F;ZNVuIU6UUiM" rel="nofollow">https:&#x2F;&#x2F;youtu.be&#x2F;ZNVuIU6UUiM ).

        The books are there.

        What&#x27;s happening is that these books are getting bought by AI companies and scanned. If they weren&#x27;t bought by the AI company then... the book in the... I&#x27;m not even sure. It&#x27;s in Portland... if inventory doesn&#x27;t sell, it may instead get thrown out. Maybe this year, maybe in ten years - it costs money to have inventory that doesn&#x27;t sell.

        Libraries will do book sales of books that aren&#x27;t checked out frequently. <a href="https:&#x2F;&#x2F;www.booksalefinder.com" rel="nofollow">https:&#x2F;&#x2F;www.booksalefinder.com . Sometimes they&#x27;re donated, sometimes it&#x27;s library discards.

        What do libraries do as they cull their collections?

        <a href="https:&#x2F;&#x2F;nwls.wislib.org&#x2F;what-to-do-with-discarded-books&#x2F;" rel="nofollow">https:&#x2F;&#x2F;nwls.wislib.org&#x2F;what-to-do-with-discarded-books&#x2F;

        &gt; Some libraries will donate weeded materials to community resale shops, Goodwill, or other resale shops. Some will take all of your unwanted books without question, but many resale shops don’t have the capacity to accept the large number of books that libraries discard. If you do find a place to take them, you’ll still need to use staff or volunteer time to get them there.

        &gt; You’ll often hear this idea from well-meaning community members, and there are some fun crafts that can be made from discarded books. A quick online search will turn up many ideas for ways to use your books in crafts for kids, teens, and adults, including folded book art, blackout poetry, collage or prints made on book pages, wreaths and garlands made from book pages, and more. But even if you do a LOT of craft activities at your library, it’s unlikely that you could use up enough of your discards to even notice the difference.

        For example... <a href="https:&#x2F;&#x2F;reallifeartist.wordpress.com&#x2F;tag&#x2F;book-carving&#x2F;" rel="nofollow">https:&#x2F;&#x2F;reallifeartist.wordpress.com&#x2F;tag&#x2F;book-carving&#x2F; or <a href="https:&#x2F;&#x2F;www.cutandfoldbookart.com&#x2F;cut-and-fold-method-instructions&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.cutandfoldbookart.com&#x2F;cut-and-fold-method-instru...

        These are books that used book stores (and libraries) are trying to get rid of to free up room for inventory that will move or shelf space for books that people will check out and read.

        ... but the last copy of a book can always be found in the national library where it was published... though if you really want to preserve Genetics, Radiobiology and Radiology Proceedings, Mid-Western Conference from 1959 from being slurped up by LLM training, you can buy a copy and shelf it yourself.

        1. aprilthird2021 · · focus · HN ↗
          This isn&#x27;t true. A lot of them are buying rare books or the only known copies of books, scanning them, then pulping the hard copy
  4. t1234s · · focus · HN ↗
    The value in these AI companies will be more in their proprietary training data than the models.
  5. RobotToaster · · focus · HN ↗
    The sad part is, imagine how positive this could be if the scans were made available to the public.
    1. IrishTechie · · focus · HN ↗
      Might add insult to injury for the publishers&#x2F;authors though?
      1. doublerabbit · · focus · HN ↗
        How? A book, song that has been out in public-domain for more than 10 years, should be downloadable. Even if you were to rebuy the same CD again the amounts that the artist received would be a pittance. Why not just let it be free to be enjoyed by all?

        OCR Scanned for training, then tossed away or burnt. Great for nature.

        1. ChickeNES · · focus · HN ↗
          Why would they be tossed or burnt? Tons of discarded books are simply recycled like any other paper object (that&#x27;s why people keep saying they were &quot;pulped&quot;)
      2. RobotToaster · · focus · HN ↗
        Maybe for current in print books, but the concern is about rare books being pulped in this process. If a book is rare then it isn&#x27;t in print, so nobody is making money out of it.
        1. nightpool · · focus · HN ↗
          Unfortunately copyright law does not have a squatters-rights exception
          1. ChickeNES · · focus · HN ↗
            I&#x27;ve believed for years that it really should have one, or at least a &quot;you aren&#x27;t selling this to the public at a reasonable market-rate price (or via subscription, I care about access far FAR more than ownership), you lose all rights to it&quot; regime.
  6. genxy · · focus · HN ↗
    It doesn&#x27;t matter what language the tokens are in, now Eye of Sauron seeks to consume all knowledge.
    1. pfdietz · · focus · HN ↗
      How does one &quot;consume knowledge&quot;?
      1. wccrawford · · focus · HN ↗
        Well, in this case, I imagine they mean by shredding books after scanning them. Since that&#x27;s what&#x27;s happening.
        1. pfdietz · · focus · HN ↗
          That&#x27;s not consuming knowledge, that&#x27;s consuming cellulose and ink.
          1. genxy · · focus · HN ↗
            When the refcount goes to zero the knowledge is consumed. Your take is overlay pedantic without benefit.

            Both definitions of knowledge creation can used. If one creates a book and it is never read, has it been produced? Isn&#x27;t knowledge also its access and how widely it is disseminated?

            1. pfdietz · · focus · HN ↗
              &gt; Your take is overlay pedantic without benefit.

              The benefit is that it properly mocks a misleading framing.

              &gt; If one creates a book and it is never read, has it been produced?

              If one destructively scans a book that will never be read, has anything been lost?

              1. genxy · · focus · HN ↗
                Then say what you mean, instead of having coy false confusion and this addled manic line of conversation.

                Very little was gained.

  7. rjh29 · · focus · HN ↗
    Good memories of visiting used bookshops near Kyoto University with stacks upon stacks of obscure literary works and research material. Lots of interesting books about the Japanese language that were never digitized and I always left with 2-3 new books. So I&#x27;m not super happy about AI companies hoovering this all up and not making the scans available.
  8. applfanboysbgon · · focus · HN ↗
    At this point I&#x27;m starting to think venture capital has got to go. What an unbelievably destructive ideology. (Yes, I&#x27;m aware I&#x27;m posting on VCnews)
  9. zirkonit · · focus · HN ↗
    Our family volunteers at a nonprofit that moves a huge number of books. We take donations and run massive charity sales, clearing tens of thousands of books a month. Pricing works like a ladder: you try to sell a book for a couple of bucks, then for a dollar, then by the $5 bag, then for free, and you still end up with thousands of books nobody wants even at no cost. These used to go straight to pulp. Now they go to AI labs for scanning. Would we rather they were read, or at least owned, by someone? Yes. Is scanning better than turning them into toilet paper? Yes, even if only marginally.

    I&#x27;m sure the Anthropics of the world have shredded the last known copies of plenty of titles. But nobody talks about the last known copies that get pulped every day simply because there are so many books nobody wants and storing them is expensive.

    1. pixl97 · · focus · HN ↗
      Yea, there are a ton of people that seem they&#x27;d rather the books get lost forever than be looked at by an AI company.
    2. szszrk · · focus · HN ↗
      My wife loves bulk book hauls. Those places that we frequent, work like a literal permanent discount warehouse - books in high shelves, on pallets, everywhere. Often hundreds of issues of the same one.

      But there is so many books there that no one want&#x27;s to read. Hundreds of the same book lying there for months or years.

      Same for public book-sharing &quot;libraries&quot; (small shelves that look like bird house, usually in parks etc). People really like them and there are many in my city, but most books there are products of a gone era and a gone mindset. No one want&#x27;s that even for free.

      We were taught respect for books, but not everything is worth preserving.

      1. ghaff · · focus · HN ↗
        I didn&#x27;t go this year but my local town library has a book sale where books go for something like $10&#x2F;bag. I donate some books to them throughout the year. There are still a lot of books available on the last or second to last day of the sale. I&#x27;m sure a huge number get pulped.
      2. fakedang · · focus · HN ↗
        A few days back, I spent some time going through a World Book encyclopaedia set that my parents had bought for me when I was 4. It was basically an expensive investment back then into what would turn into a lifelong reading habit.

        Most of the information in it is practically worthless today unfortunately - Saddam Hussein still alive, as were Katharine Hepburn and King Hussein of Jordan, no mention of Ceres and Pluto was still a planet, and the Internet and software had barely a mention. Most articles had less information than a standard Wikipedia article.

        That being said, the way information was presented in them still outshines anything one may find on the internet today. Even the simple elements - neat diagrams and relevant images, proper sectioning and organization of the text, footnotes to other relevant articles...

        Long gone are the days when I would simply take a bowl of ice cream and a volume and just read it end to end.

    3. sly010 · · focus · HN ↗
      Well, they could turn the bad faith story into a good faith story by making them available for everyone to download perhaps. (AI companies &quot;saving&quot; old books!) But that would require giving a s*t which they don&#x27;t and that is the real problem imho.
      1. Aurornis · · focus · HN ↗
        They legally cannot do this.
        1. theroadnotbacon · · focus · HN ↗
          That certainly hasn’t stopped them before… IP theft is kind of their whole thing, isn’t it?
          1. Aurornis · · focus · HN ↗
            Distributing copyrighted works (prior to expiration of their copyright) verbatim is illegal.

            Training an LLM on copyrighted works is not illegal.

            This whole debate has been tried in court already. Calling it IP theft only stands on individual moral grounds, but the law allows for derivative works.

          2. syrrim · · focus · HN ↗
            Google attempted to do this 15 years ago, they got sued and stopped. It turns out that tech companies occasionally do have to follow the law, you&#x27;d think people would be happier about that...
            1. keeda · · focus · HN ↗
              Wait, if you’re talking about Google Books case, Google won. Maybe they made adjustments on how they served results but they certainly did not stop.
              1. ndiddy · · focus · HN ↗
                The Google Books settlement was originally going to make Google into a clearinghouse for scans of out-of-print books. The scans would have been available for individuals to purchase for a reasonable price, and libraries and institutions would have been able to subscribe to a service that would give patrons access to the full text of every book. This deal fell apart because some research libraries and authors argued this was anti-competitive, as anyone wanting to make a competing service would have to go through the same process as Google of settling a class action lawsuit. They instead wanted Congress to pass a law to free up the rights to orphaned books. No progress on this was subsequently made because nobody in Congress cares enough about the rights to out-of-print books to get legislation passed. The whole reason why they&#x27;re out of print when ebooks and print-on-demand exist is that they won&#x27;t get enough sales to make it worth the time and money to figure out who the royalties should go to. The result is that nobody outside Google gets to see the full Google Books scans.
                1. shagie · · focus · HN ↗
                  Those scans are held at HathiTrust Research Center <a href="https:&#x2F;&#x2F;www.hathitrust.org&#x2F;about&#x2F;research-center&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.hathitrust.org&#x2F;about&#x2F;research-center&#x2F; <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;HathiTrust" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;HathiTrust

                  &gt; HathiTrust Digital Library is a large-scale collaborative repository of digital content from research libraries, administered by the University of Michigan. Its holdings include content digitized via Google Books and the Internet Archive digitization initiatives, as well as content digitized locally by libraries.

                  <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Authors_Guild,_Inc._v._HathiTrust" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Authors_Guild,_Inc._v._HathiTr... and <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Authors_Guild,_Inc._v._Google,_Inc.#Impact" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Authors_Guild,_Inc._v._Google,...

                  &gt; Authors Guild, Inc. v. HathiTrust (2014) was a following case related to HathiTrust, a project by the libraries of the Big Ten Academic Alliance and the University of California systems that combined their digital library collections with those of Google&#x27;s Book Search. The HathiTrust case differed in two primary factors which were raised by the plaintiffs: that for viewers with disabilities, they could view the scanned text through a screen reader to make it easier to read, and offering to print out the scans as replacement copies for members of the universities if they could verify their original copies were lost or damaged. Both uses were deemed also to be fair use by the Second Circuit.

                  &gt; The subject of the copyright of orphan works – works that may still be under copyright but with no identifiable rights holder – was a significant point of debate after both this and HathiTrust. Normally, libraries have been hesitant to loan digital copies of orphaned works as libraries may be liable for copyright violations should the copyright owner step forward to claim ownership.

                  The bill on orphan works that didn&#x27;t pass was <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Shawn_Bentley_Orphan_Works_Act_of_2008" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Shawn_Bentley_Orphan_Works_Act...

                  <a href="https:&#x2F;&#x2F;www.hathitrust.org&#x2F;the-collection&#x2F;search-access&#x2F;copyright-access&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.hathitrust.org&#x2F;the-collection&#x2F;search-access&#x2F;copy...

                  And there are exceptions for copyrighted works allowing them to lend them out.

                  &gt; Protected by copyright law, but made available: Protected by copyright law but made available on a strictly limited basis in accordance with the statutory limitations including, but not limited to, Section 107 provisions for fair use, Section 108 provisions for libraries and archives, and the rights provided to registered users with disabilities. In the absence of an applicable exception, no further reproduction or distribution is permitted by any means without the permission of the copyright holder. Lawful uses of works are provided only under the following conditions ...

                  Scanning still continues. <a href="https:&#x2F;&#x2F;www.hathitrust.org&#x2F;member-libraries&#x2F;contribute-content&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.hathitrust.org&#x2F;member-libraries&#x2F;contribute-conte... - though it&#x27;s not at the same rate as it was during google books project.

        2. sly010 · · focus · HN ↗
          That would require effort (to sort, acquire copyright, etc) which they wouldn&#x27;t put in. Because they don&#x27;t care.

          People obviously feel bad about companies doing this. People reading these stories don&#x27;t care what&#x27;s legal, they care what&#x27;s ethical. Heck, re-publishing long lost material would make AI companies heroes instead of bad guys.

          1. Aurornis · · focus · HN ↗
            &gt; That would require effort (to sort, acquire copyright, etc) which they wouldn&#x27;t put in. Because they don&#x27;t care.

            I don&#x27;t think you have any idea how expensive it is to acquire the copyright for a single book with the intent of making it freely available online. That&#x27;s equivalent to asking the rights holders to perpetually forgo all possible earnings from the material, and they expect to be compensated accordingly. Even paying lawyers to begin assembling what&#x27;s needed to make this happen would be five figures per book to get started.

        3. 2OEH8eoCRo0 · · focus · HN ↗
          They legally cannot scan them in entirety either but they are.
          1. ChickeNES · · focus · HN ↗
            Again, under Bartz v Anthropic they can scan and train on whatever they want, as long as the original is lost in the process.
            1. 2OEH8eoCRo0 · · focus · HN ↗
              TIL thanks.
        4. wareya · · focus · HN ↗
          What I would do is announce that the scans are being preserved and will donated to the library of congress or whatever other institution is legally able to hold onto stuff like this. But that would more directly tie specific companies to the practice of destroying unwanted old books, so nobody&#x27;s going to do it, or even announce it.
          1. shagie · · focus · HN ↗
            [delayed]
            1. wareya · · focus · HN ↗
              The library of congress does not have copies of random japanese books.
              1. shagie · · focus · HN ↗
                [delayed]
      2. nightpool · · focus · HN ↗
        Both Anthropic and Internet Archive have gotten sued into oblivion by publishing companies once, there&#x27;s no way they&#x27;re going to do something even more flagrantly illegal
    4. icantevenhold · · focus · HN ↗
      Why do they shred them at all instead of donating or selling them again?
      1. quickthrowman · · focus · HN ↗
        A judge ruled it was OK to scan books and save the scanned copy if you shred the physical book afterwards.
        1. shagie · · focus · HN ↗
          [delayed]
      2. andrew_lettuce · · focus · HN ↗
        My understanding is they cut off the binding for scanning. They&#x27;d have to resell it donate by the page
        1. ChickeNES · · focus · HN ↗
          They have to destroy the copy either way for it to be fair use (according to Bartz v Anthropic). Cutting the bindings off is already the faster method to scan, and once you&#x27;re required to pulp the original anyway, it becomes a no-brainer.
          1. shagie · · focus · HN ↗
            [delayed]
    5. laybak · · focus · HN ↗
      I have a similar thought too each time I walk past piles of discount books.

      I&#x27;m in the camp that perhaps it&#x27;s healthy to not grasp onto every bit of information. that some artifacts dying a natural death is maybe just the way things are

    6. palmotea · · focus · HN ↗
      &gt; Now they go to AI labs for scanning. Would we rather they were read, or at least owned, by someone? Yes. Is scanning better than turning them into toilet paper? Yes, even if only marginally.

      Yeah, now instead of old books being turned into toilet paper, we&#x27;ll get turned into toilet paper. Much better.

      But Sam Altman will become richer than God, and isn&#x27;t that what really matters?

      But don&#x27;t worry! You&#x27;ll still have access to ChatGPT until your savings run out.

      1. ChickeNES · · focus · HN ↗
        Do you have an actual point about scanning the books, or are you just using this as a soapbox to rant about AI and Sam Altman?
        1. palmotea · · focus · HN ↗
          Yes, you missed it. Perhaps you should read it again until you get it?
          1. ChickeNES · · focus · HN ↗
            Again, do you have an actual objection to the subject of the article, or are you just mad it enriches people you don&#x27;t like?
            1. palmotea · · focus · HN ↗
              &gt; Again, do you have an actual objection to the subject of the article

              I was responding to a comment, which apparently is another thing you missed.

              &gt; or are you just mad it enriches people you don&#x27;t like?

              Nice strawman, pity if someone knocked it down.

              1. ChickeNES · · focus · HN ↗
                Okay, so all you have are personal attacks, goodbye.
                1. palmotea · · focus · HN ↗
                  Note: I made no personal attacks.
    7. a_shovel · · focus · HN ↗
      A lot of these books are reference titles that are outdated to the point of uselessness and&#x2F;or weren&#x27;t that great&#x2F;interesting when they were new. Not many people have interest in or use for textbooks from the 50s.
  10. tsylba · · focus · HN ↗
    Ah yes, the litteral destruction of culture and physical media for a centralised subscription service. I love the liberal world of techno enclosures of our new overlords, viva el free market economy.
    1. Analemma_ · · focus · HN ↗
      What do you think happened to all these used books before the AI companies showed up?
      1. alightsoul · · focus · HN ↗
        They sat in boxes
        1. ghc · · focus · HN ↗
          Not quite:

          &gt; But what happens when sales numbers don&#x27;t meet projections? The book is discounted. Then, at the publisher&#x27;s discretion, the bookstore will receive a directive to rip the covers off the books, recycle the remainder of the book to be &quot;pulped&quot; or turned into other forms of paper, such as notebook paper and toilet paper. The bookstore is expected to mail the book covers to the publisher as evidence that the book has been destroyed.

          <a href="https:&#x2F;&#x2F;www.offthebeatenshelf.com&#x2F;blog&#x2F;pulp-fiction-is-real" rel="nofollow">https:&#x2F;&#x2F;www.offthebeatenshelf.com&#x2F;blog&#x2F;pulp-fiction-is-real

          1. sebmellen · · focus · HN ↗
            TIL that&#x27;s where the term &quot;pulp fiction&quot; comes from!
          2. alightsoul · · focus · HN ↗
            That&#x27;s an American take. In Latin America they sit in boxes and you can see it, because the pages will have been discolored after so many years of sitting on a shelf waiting to be sold
            1. ChickeNES · · focus · HN ↗
              Where in Latin America?

              <a href="https:&#x2F;&#x2F;www1.folha.uol.com.br&#x2F;ilustrada&#x2F;2016&#x2F;09&#x2F;1818064-se-nao-forem-vendidos-ou-doados-livros-podem-virar-papel-higienico.shtml" rel="nofollow">https:&#x2F;&#x2F;www1.folha.uol.com.br&#x2F;ilustrada&#x2F;2016&#x2F;09&#x2F;1818064-se-n...

              <a href="https:&#x2F;&#x2F;www.jornada.com.mx&#x2F;2005&#x2F;05&#x2F;25&#x2F;index.php?article=a04n1cul&amp;section=cultura" rel="nofollow">https:&#x2F;&#x2F;www.jornada.com.mx&#x2F;2005&#x2F;05&#x2F;25&#x2F;index.php?article=a04n...

              <a href="https:&#x2F;&#x2F;www.lanacion.com.ar&#x2F;cultura&#x2F;a-donde-van-libros-no-venden-entre-nid2172443&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.lanacion.com.ar&#x2F;cultura&#x2F;a-donde-van-libros-no-ve...

              <a href="https:&#x2F;&#x2F;revistadossier.udp.cl&#x2F;dossier&#x2F;destruir-con-permiso&#x2F;" rel="nofollow">https:&#x2F;&#x2F;revistadossier.udp.cl&#x2F;dossier&#x2F;destruir-con-permiso&#x2F;

              Looks to me, going by this reporting, that LATAM treats books identically to the US (The first headline, translated, is literally &quot;If not sold or donated, books can become toilet paper&quot;)

              1. alightsoul · · focus · HN ↗
                In central America for example I went to Panama&#x27;s book fair and you could see old books for sale. There&#x27;s a lot of fear among booksellers that a book will not sell, because they purely resell them so any losses fall entirely on booksellers. They do not have any partnerships with publishers
        2. scottyah · · focus · HN ↗
          That only lasts 10-30yrs. They&#x27;re cheap paper, they were never made to last.
      2. ghc · · focus · HN ↗
        &quot;Once great literature—now great litter.&quot;
    2. quickthrowman · · focus · HN ↗
      Feel free to buy books by the ton and preserve them yourself. The simple fact they’re being sold by weight implies they’re not rare or unique.
      1. mistercheph · · focus · HN ↗
        The simple fact that the AI labs are spending billions of dollars to acquire and scan them implies they are rare and unique.
        1. quickthrowman · · focus · HN ↗
          Please provide proof for the billions of dollars claim, thank you.
          1. mistercheph · · focus · HN ↗
            25 million books

            scanning+ocr: 0.05 per page * 300 pgs: 15&#x2F;bk

            acquisition+disposal cost: 2.00&#x2F;bk

            total acq+scan+dispose = 425M

            settlement value &#x2F; bk: $3,000 [bartz v anthropic]

            settlement probability, let&#x27;s say 1%: 0.01

            expected settlement &#x2F; bk : $30

            expected settlement total: 750M [industry-wide this is certainly an underestimate]

            legal fees: 25% of settlement total: 187M

            total 1.36B

            + operations, engineering, storage

  11. fangspire · · focus · HN ↗

    [dead]

  12. patall · · focus · HN ↗
    Can someone explain why the old books could really be relevant. I get the pre-nuclear steel analogy, but why is this relevant given how much more modern texts exist. A few years ago, millions of yahoo groups were erased but now a few thousand books are what is needed to run a successful AI company? I mean, it can barely be about the information in those books (that would be very often outdated), but just for a little more text (with ever less marginal gain), what is the benefit?
    1. nemomarx · · focus · HN ↗
      Writing style maybe?
    2. layer8 · · focus · HN ↗
      Diversity. Modern books in modern styles are overrepresented, old ones underrepresented.
    3. aisenik · · focus · HN ↗
      Cognition is encoded in language, they weren&#x27;t brain-damaged yet. There&#x27;s better (real, not token-exchange) thinking, which LLMs can copy and reproduce in novel arrangements.
    4. mars-or_wars · · focus · HN ↗

      [dead]

    5. aprilthird2021 · · focus · HN ↗
      The information density is a lot less for millions of yahoo groups. They are far more likely to cover the same topics and not have new information in them.

      Books are more likely to be about a specific topic or story or time or setting and be more information dense

    6. hyperhello · · focus · HN ↗
      They are the memories of the productive part of society. You can leaf through them and get the feel of what it was like. You don’t need most of your memories, personality, or core ideals to be productive to the State.
    7. Legend2440 · · focus · HN ↗
      It sounds like it is about the information in the books. The titles they&#x27;re looking for are all nonfiction. Not everything is on the internet, and just because it&#x27;s a few years old doesn&#x27;t mean it&#x27;s outdated.

      Speaking from experience, the information density of published books is a lot higher than most internet text. It&#x27;s very high quality training data.

      The goal here is to have all human knowledge in a single file, which is pretty neat IMO.

      1. ChickeNES · · focus · HN ↗
        &gt; the information density of published books is a lot higher than most internet text

        I&#x27;m not convinced. I think you are under-weighing the massive volumes of stuff like self-help books, romance novels, etc.

        1. criemen · · focus · HN ↗
          They&#x27;re targeting non-fiction books, so romance novels would be out.
          1. scottyah · · focus · HN ↗
            Is that a policy change after o4 got a little out of hand?
          2. ChickeNES · · focus · HN ↗
            Booksellers noticed a huge uptick in non-fiction purchases, that is not the same thing as them not targeting fiction at all.

            Edit: Actually, I have real evidence, the Bartz in &quot;Bartz v Anthropic&quot; is Andrea Bartz, a novelist, and the complaint specifically lists four of her novels as infringed works.

      2. newsy-combi · · focus · HN ↗
        The internet basically never delivered on the promise of replacing textbooks or even education as a whole. Wikipedia sucks on many topics, has insane internal politics, and is a tertiary source by design (redigesting blogs and books), whereas textbooks are generally secondary.
      3. zardo · · focus · HN ↗
        Also the ability to set the training input limit in the past could be useful.
      4. patall · · focus · HN ↗
        I see. So it is much less about 150 year old fiction books, but more about 1980s science literature that was only ever printed five times. That makes a lot more sense than what the public debate seems to be about.
    8. timcobb · · focus · HN ↗
      I&#x27;m guessing the more integration tables they consume, the better they become at at integration. I&#x27;m guessing these companies are looking for all kinds of stuff, not just prose.

      I also guess that they&#x27;re targeting languages that aren&#x27;t tier one for them yet. Like, Japanese is probably a relatively small corpus for them.

    9. Aurornis · · focus · HN ↗
      &gt; but now a few thousand books are what is needed to run a successful AI company?

      They&#x27;re scanning millions of books.

      It&#x27;s the diversity of text that helps. One of the lessons we&#x27;ve learned is that more training data leads to better models. Even old books have different mixes of word sequences that will improve the model. The returns are diminishing, but when you have the pipeline set up to ingest it you might as well keep adding to the dataset.

    10. theroadnotbacon · · focus · HN ↗
      I also wonder if it’s used for text generation in image models! Awful lot of typefaces, sizes, orientations, and words in those books.
    11. mistercheph · · focus · HN ↗
      &gt; why the old books could really be relevant &gt; it can barely be about the information in those books (that would be very often outdated),

      LOL, not being rude: have you ever read a book outside of what they forced you to read in school? Most old books are not O&#x27;reilly&#x27;s manuals for Visual Studio 2014, they don&#x27;t go out of date.

      They are interesting to human beings for the same reason they are interesting to the labs. If it was just about quantity of text then the labs could generate text with the prev. gen model and use that alone to scale to the next model, there is something of immeasurable value contained in books (hint: it starts with an i and rhymes with bin formation).

    12. ijk · · focus · HN ↗
      Because Yahoo in particular was very good at destroying goldmines shortly before they became ultra valuable.

      There&#x27;s a number of places that destroyed vast amounts of data in the wind down of ZIRP that probably regret it now.

    13. 2OEH8eoCRo0 · · focus · HN ↗
      They aren&#x27;t but these companies have more money than sense.
    14. cesarvarela · · focus · HN ↗
      I think a book is like one completion of the mind behind it, so in a way, this is just distillation.
    15. acuozzo · · focus · HN ↗
      &gt; Can someone explain why the old books could really be relevant.

      Lots and lots of information is not online. You&#x27;d be surprised.

    16. lofaszvanitt · · focus · HN ↗
      Well, most of the old books are much easier to read than the new ones. Back then people were much more educated and had a much wider vocabulary....
  13. YVoyiatzis · · focus · HN ↗
    Pretty much as it happened with vinyl records twenty years ago. I remember seeing photos of this guy somewhere in Brazil standing atop heaps of vinyl records which he had amassed with HDLR intention. Now books. Some of us hold on forever.
    1. mistercheph · · focus · HN ↗
      You are a moron: vinyl records came and went in about 25 years, books have been the engine of human progress for at least the last 3 millenia, the books being burned by these misanthropic lunatics are not available in any other medium, this is not about fascination with some particular mediumn of transmission, it&#x27;s about the contents
      1. ChickeNES · · focus · HN ↗
        &quot;the engine of human progress&quot;? I think you mean capitalism.

        &gt; the books being burned by these misanthropic lunatics are not available in any other medium

        prove it, name one title

        &gt; this is not about fascination with some particular mediumn of transmission

        it very much is. this fetishism of books should really stop, especially when ebooks are more useful, durable, etc.

        1. mistercheph · · focus · HN ↗
          &gt; prove it, name one title

          if they were available, the labs would have purchased them digitally or pirated them

  14. mistercheph · · focus · HN ↗
    Before you buy the apologia that these books are not valuable or interesting: if they weren&#x27;t valuable or interesting the AI labs would not be spending billions of dollars to purchase and scan them. Yes, everyone has a personal anecdote about pallets of garbage books but I have three insights for you that you may not have because you don&#x27;t read or sift through pallets of garbage books:

    1) The labs don&#x27;t want garbage books, they want interesting books that are rare and unique. They want high quality training data, random permutations of language style are fine, but what you want is unseen information, unseen patterns of thinking, unseen ideas.

    2) Most pallet of books contains lots of valuable and interesting works, maybe 1-3% but sorting through them takes time, money, and energy, that&#x27;s why the labs are starting to purchase by the pallet, it&#x27;s because they already have a fully automated process so they can always beat any bookseller small or large on cost to find the books of interest and value in a pile.

    3) Many of these pallets may sit for years before being sorted, and many of the books may sit for years before being sold, but these things actually do eventually happen, valuable books are found, and they eventually make their way to interested readers, this is the business model of used bookstores. Most books of value don&#x27;t get destroyed or thrown away.

    Destroying human art, knowledge, and culture is an essential part of the business plan for frontier labs, it is not enough to steal and regurgitate all the art and information in the world, you also want to make it inaccessible through any other means than the regurgitation machine. Don&#x27;t expect the book burning to be an isolated incident, they are coming for every other form of stored human knowledge or art, and yes, unfortunately while scanning it they will have to destroy the original copy. And attacking the past is only the beginning.

    1. ChickeNES · · focus · HN ↗
      I have sifted through everything from piled up junky independent book shops, to cast offs from research libraries, to dumpsters of end of life books. Most books really are not worth the paper they are printed on.

      &gt; Most books of value don&#x27;t get destroyed or thrown away.

      Of value to who? Most used bookstores are boutiques that over-curate and will happily refuse or recycle books that they deem are inferior&#x2F;irrelevant. This is a big reason why I prefer Half Price Books over most any other used bookstore, they sell most everything.

      1. mistercheph · · focus · HN ↗
        &gt; Most books really are not worth the paper they are printed on.

        Yeah, but like I said in the comment you are replying to, somewhere around 1-5% of books are worth the paper they are printed on and far more, and those are exactly the books the labs are trying to scan and destroy, they are not buying pallets of books and scanning them in order to scan the june 1995 tv guide for the 800th time, they want unique, interesting, and high quality text, because that&#x27;s what feeds pretraining.

    2. x-complexity · · focus · HN ↗
      (3) is doing almost all of the moral lifting, in that it assumes that relatively few of the books get pulped in the first place:

      &gt; 3) Many of these pallets may sit for years before being sorted, and many of the books may sit for years before being sold, but these things actually do eventually happen, valuable books are found, and they eventually make their way to interested readers, this is the business model of used bookstores. Most books of value don&#x27;t get destroyed or thrown away.

      This contradicts base reality, as the vast majority of the books in such cases were already queued to be pulped.

      &gt; Destroying human art, knowledge, and culture is an essential part of the business plan for frontier labs, it is not enough to steal and regurgitate all the art and information in the world, you also want to make it inaccessible through any other means than the regurgitation machine.

      Even if Anthropic wanted to make the books publicly available as a form of moralwashing, they legally can&#x27;t do that. Their current process rides the line of fair use, in that it obeys the requirement of medium shifting to be considered legally defensible.

      Don&#x27;t like the outcome? Change the law to allow for digital archival without requiring destruction.

  15. dofm · · focus · HN ↗
    AI firms buying and destroying the sources of knowledge is a weird way to get us to Fahrenheit 451 but maybe the USA will get there before it switches to metric after all.
  16. offsky · · focus · HN ↗
    Business plan: Use AI to write fake old books. Make a used bookstore to sell these books back to AI for training.
    1. financetechbro · · focus · HN ↗
      Circular economy maxxxing
  17. andrekandre · · focus · HN ↗
    once these books have all been scanned and destroyed, doesn&#x27;t it mean later competitors (and free-as-in-foss alternatives) wont ever be able to catch up?

    it seems really sad that these books data isnt open in the first place...

  18. Theory42 · · focus · HN ↗
    You&#x27;d think they could give the scans to archive.org or similar. Even if the books are in copyright, they won&#x27;t be forever. We can&#x27;t let our culture be destroyed like this.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.