‹ BackHN Continuity

Thread

Japanese used bookstores see 5x sales surge as books are being bought by the ton

96 points · 130 comments · speckx

  1. patall · · focus · HN ↗
    Can someone explain why the old books could really be relevant. I get the pre-nuclear steel analogy, but why is this relevant given how much more modern texts exist. A few years ago, millions of yahoo groups were erased but now a few thousand books are what is needed to run a successful AI company? I mean, it can barely be about the information in those books (that would be very often outdated), but just for a little more text (with ever less marginal gain), what is the benefit?
    1. nemomarx · · focus · HN ↗
      Writing style maybe?
    2. layer8 · · focus · HN ↗
      Diversity. Modern books with modern content and in modern styles are overrepresented, old ones underrepresented.
    3. aisenik · · focus · HN ↗
      Cognition is encoded in language, they weren't brain-damaged yet. There's better (real, not token-exchange) thinking, which LLMs can copy and reproduce in novel arrangements.
    4. mars-or_wars · · focus · HN ↗

      [dead]

    5. aprilthird2021 · · focus · HN ↗
      The information density is a lot less for millions of yahoo groups. They are far more likely to cover the same topics and not have new information in them.

      Books are more likely to be about a specific topic or story or time or setting and be more information dense

    6. hyperhello · · focus · HN ↗
      They are the memories of the productive part of society. You can leaf through them and get the feel of what it was like. You don’t need most of your memories, personality, or core ideals to be productive to the State.
    7. Legend2440 · · focus · HN ↗
      It sounds like it is about the information in the books. The titles they're looking for are all nonfiction. Not everything is on the internet, and just because it's a few years old doesn't mean it's outdated.

      Speaking from experience, the information density of published books is a lot higher than most internet text. It's very high quality training data.

      The goal here is to have all human knowledge in a single file, which is pretty neat IMO.

      1. ChickeNES · · focus · HN ↗
        > the information density of published books is a lot higher than most internet text

        I'm not convinced. I think you are under-weighing the massive volumes of stuff like self-help books, romance novels, etc.

        1. criemen · · focus · HN ↗
          They're targeting non-fiction books, so romance novels would be out.
          1. scottyah · · focus · HN ↗
            Is that a policy change after o4 got a little out of hand?
          2. ChickeNES · · focus · HN ↗
            Booksellers noticed a huge uptick in non-fiction purchases, that is not the same thing as them not targeting fiction at all.

            Edit: Actually, I have real evidence, the Bartz in "Bartz v Anthropic" is Andrea Bartz, a novelist, and the complaint specifically lists four of her novels as infringed works.

      2. newsy-combi · · focus · HN ↗
        The internet basically never delivered on the promise of replacing textbooks or even education as a whole. Wikipedia sucks on many topics, has insane internal politics, and is a tertiary source by design (redigesting blogs and books), whereas textbooks are generally secondary.
      3. zardo · · focus · HN ↗
        Also the ability to set the training input limit in the past could be useful.
      4. patall · · focus · HN ↗
        I see. So it is much less about 150 year old fiction books, but more about 1980s science literature that was only ever printed five times. That makes a lot more sense than what the public debate seems to be about.
    8. timcobb · · focus · HN ↗
      I'm guessing the more integration tables they consume, the better they become at integration.

      I also guess that they're targeting languages that aren't tier one for them yet. Like, Japanese is probably a relatively small corpus for them.

    9. Aurornis · · focus · HN ↗
      > but now a few thousand books are what is needed to run a successful AI company?

      They're scanning millions of books.

      It's the diversity of text that helps. One of the lessons we've learned is that more training data leads to better models. Even old books have different mixes of word sequences that will improve the model. The returns are diminishing, but when you have the pipeline set up to ingest it you might as well keep adding to the dataset.

    10. theroadnotbacon · · focus · HN ↗
      I also wonder if it’s used for text generation in image models! Awful lot of typefaces, sizes, orientations, and words in those books.
    11. mistercheph · · focus · HN ↗
      > why the old books could really be relevant > it can barely be about the information in those books (that would be very often outdated),

      LOL, not being rude: have you ever read a book outside of what they forced you to read in school? Most old books are not O'reilly's manuals for Visual Studio 2014, they don't go out of date.

      They are interesting to human beings for the same reason they are interesting to the labs. If it was just about quantity of text then the labs could generate text with the prev. gen model and use that alone to scale to the next model, there is something of immeasurable value contained in books (hint: it starts with an i and rhymes with bin formation).

    12. ijk · · focus · HN ↗
      Because Yahoo in particular was very good at destroying goldmines shortly before they became ultra valuable.

      There's a number of places that destroyed vast amounts of data in the wind down of ZIRP that probably regret it now.

    13. 2OEH8eoCRo0 · · focus · HN ↗
      They aren't but these companies have more money than sense.
    14. cesarvarela · · focus · HN ↗
      I think a book is like one completion of the mind behind it, so in a way, this is just distillation.
    15. acuozzo · · focus · HN ↗
      > Can someone explain why the old books could really be relevant.

      Lots and lots of information is not online. You'd be surprised.

    16. lofaszvanitt · · focus · HN ↗
      Well, most of the old books are much easier to read than the new ones. Back then people were much more educated and had a much wider vocabulary....
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.