‹ BackHN Continuity

Thread

Unsealed Briefs in Authors’ Case v. Microsoft/OpenAI

631 points · 623 comments · papergirl

  1. Skyy93 · · focus · HN ↗
    This is a lobby organisation using only the pieces and bits they like to push their own agenda.

    "sketchy russian website", how about using some more clear description like: A library for sharing books and articles that should be partly public domain because they were paid for by the public. Only some of the material is copyrighted by authors. However, some of their work is so old that it is not reprinted anyway.

    But of course such an explanation would not click.

    I also don't see a problem with statements about making people jobless. Imagine if every robotic or automation company advertised like this: Yeah, you'll buy tons of expensive robots and still rely on expensive labor from real people without any efficiency gains.

    1. trompetenaccoun · · focus · HN ↗
      It can be a bit confusing due to the terrible style of the article (ironic given the source) but it seems the "sketchy russian website" part is a direct quote by Anthropic's Sam McCandlish. And apparently Dario Amodei referred to it as sketchy as well.

      I find the brazenness of saying this while running what's arguably the largest copyright theft operation in human history astonishing. If libgen is "sketchy", then what is OpenAI?

      1. qarl · · focus · HN ↗
        > the largest copyright theft operation in human history

        Many people think that it was fair use: training is akin to reading, not copying.

        Especially the courts.

        1. Trusteando · · focus · HN ↗
          Reading isn't the right comparison. Human memory is lossy and fades while LLM encoding is durable with no degradation. The valuable content that the author provides: the content, style, selection of topics, and more is encoded, written into the LLM weights, and they obtain profit from them (now directly, via ads). No one can compete with that kind of copying and pasting from copyright-protected material. And the scale is what hurts authors most: flooding the market with millions of copies on demand, without paying for it.
          1. qarl · · focus · HN ↗
            > No only reading, because the content, style, selection of topics, and more is encoded, written, in the LLM weights

            Exactly analogous to a human reading the material.

            1. asutekku · · focus · HN ↗
              Most people can't recite a book they've read verbatim. However, if you ask an LLM to continue a random sentence from a semi-popular book, it can sometimes provide the exact text, unless the system flags the response.
              1. qarl · · focus · HN ↗
                > Most people can't recite a book

                Yes - but some people can.

                Are they criminals for reading books?

                1. asutekku · · focus · HN ↗
                  Those some people can't recite a book to any given person in the world at any given time. There's a difference.
                  1. qarl · · focus · HN ↗
                    So if we invent a way for those people to talk to everyone in the world - then they would be a criminal for reading books whether they did or not?

                    You're not making sense.

                    The problem is the reciting. Not the reading. And hence, not the training.

                    1. asutekku · · focus · HN ↗
                      If you would go on a tv and read out loud a book you have not purchased rights to present, according to the most copyright laws in the world, yes you would and you would get a fine. Similarly if you broadcast a tv-show you have not purchased rights would.

                      Whether this is right or not is a seperate question.

                      1. qarl · · focus · HN ↗
                        > If you would go on a tv and read out loud a book

                        That is not the situation we are discussing. No one is arguing that what you describe is infringement.

                        What we are discussing is the training - which happens BEFORE the broadcast. It is analogous to reading. Is simply READING the material an infringement.

                        1. asutekku · · focus · HN ↗
                          Using analogies like "reading" to describe AI training is quite misleading imho. Training effectively embeds the book's contents into the model's weights. Changing the format doesn't change the content; a better analogy is distributing a book's text within software.

                          Current copyright laws are simply not prepared for this unprecedented use.

                          1. qarl · · focus · HN ↗
                            > Training effectively embeds the book's contents into the model's weights. Changing the format doesn't change the content

                            Reading a book embeds its contents into your brain. And yet, that is considered fair use.

                            I agree, the analogies are meaningless in a legal context. In the legal context, the courts disagree with you.

                            1. efreak · · focus · HN ↗
                              I'm not disagreeing with you, but the comparison is breaking down here. Reading a book is the specific intended purpose. It's why the book exists, not just fair use. If you couldn't remember what was going on in the book as you read it, it would be worthless.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.