‹ BackHN Continuity

Thread

Unsealed Briefs in Authors’ Case v. Microsoft/OpenAI

631 points · 623 comments · papergirl

  1. Skyy93 · · focus · HN ↗
    This is a lobby organisation using only the pieces and bits they like to push their own agenda.

    "sketchy russian website", how about using some more clear description like: A library for sharing books and articles that should be partly public domain because they were paid for by the public. Only some of the material is copyrighted by authors. However, some of their work is so old that it is not reprinted anyway.

    But of course such an explanation would not click.

    I also don't see a problem with statements about making people jobless. Imagine if every robotic or automation company advertised like this: Yeah, you'll buy tons of expensive robots and still rely on expensive labor from real people without any efficiency gains.

    1. trompetenaccoun · · focus · HN ↗
      It can be a bit confusing due to the terrible style of the article (ironic given the source) but it seems the "sketchy russian website" part is a direct quote by Anthropic's Sam McCandlish. And apparently Dario Amodei referred to it as sketchy as well.

      I find the brazenness of saying this while running what's arguably the largest copyright theft operation in human history astonishing. If libgen is "sketchy", then what is OpenAI?

      1. qarl · · focus · HN ↗
        > the largest copyright theft operation in human history

        Many people think that it was fair use: training is akin to reading, not copying.

        Especially the courts.

        1. trompetenaccoun · · focus · HN ↗
          That's not been legally established, the litigation is ongoing. And if mere downloading and reading of copyrighted material were legal, how come torrent users have been fined for it in the thousands?

          The law is the law, there can't be different law for corporations with billions in backing. I don't agree with current copyright laws btw and think they should be changed. However, they probably should have lobbied for that before illegally downloading all this material.

          1. qarl · · focus · HN ↗
            > That's not been legally established

            100% of the rulings agree with me.

            The piracy is not in question. It is unarguably copyright violation.

            But that's not what anyone means in this context. Training is what everyone means.

            > The law is the law, there can't be different law for corporations with billions in backing.

            I didn't say otherwise. That's a straw man.

            1. trompetenaccoun · · focus · HN ↗
              So we agree they have violated copyright at a much larger scale than LibGen, yet they call LibGen "sketchy" for doing the same thing? Absurd, what exactly are we arguing here?
              1. qarl · · focus · HN ↗
                > So we agree they have violated copyright at a much larger scale than LibGen

                No.

            2. latexr · · focus · HN ↗
              So, which is it? First you said it was fair use, now you’re saying it’s unarguably copyright violation. You can’t have it both ways.
              1. qarl · · focus · HN ↗
                Training is fair use - the torrenting was infringement. Sorry if that was confusing.
                1. mrdependable · · focus · HN ↗
                  Training has not, in fact, been decided on yet. Of the four criteria deciding whether something is fair use, training only maybe passes two.

                  <a href="https:&#x2F;&#x2F;www.copyright.gov&#x2F;ai&#x2F;Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf" rel="nofollow">https:&#x2F;&#x2F;www.copyright.gov&#x2F;ai&#x2F;Copyright-and-Artificial-Intell...

                  1. qarl · · focus · HN ↗

                    [dead]

                  2. tpmoney · · focus · HN ↗
                    Training in the US has, in fact, been decided (at least to the extent that anything has currently been decided). Bartz vs Anthropic specifically ruled training an AI model on legally owned copyrighted material is sufficiently transformative[1]:

                        This order grants summary judgment for Anthropic that the training use was a fair use.
                        And, it grants that the print-to-digital format change was a fair use for a different reason. But it
                        denies summary judgment for Anthropic that the pirated library copies must be treated as
                        training copies.
                    
                    The document you linked was written a month before the Bartz decision was reached. It&#x27;s also worth noting even the document you linked says this in its conclusion:

                        Various uses of copyrighted works in AI training are likely to be transformative. The
                        extent to which they are fair, however, will depend on what works were used, from what
                        source, for what purpose, and with what controls on the outputs—all of which can affect the
                        market.
                    
                    
                    [1]: <a href="https:&#x2F;&#x2F;copyrightalliance.org&#x2F;wp-content&#x2F;uploads&#x2F;2025&#x2F;06&#x2F;Bartz-v.-Anthropic-Order.pdf" rel="nofollow">https:&#x2F;&#x2F;copyrightalliance.org&#x2F;wp-content&#x2F;uploads&#x2F;2025&#x2F;06&#x2F;Bar...
                    1. mrdependable · · focus · HN ↗
                      That is not binding, hence the NYT trial.

                      That is not what the text you quoted is saying. It says that the output may be transformative, which is one criteria, but the other criteria depends on how it is used and what the source is.

          2. echoangle · · focus · HN ↗
            &gt; And if mere downloading and reading of copyrighted material were legal, how come torrent users have been fined for it in the thousands?

            Because torrenting includes redistribution, since you also upload to other peers

            1. wongarsu · · focus · HN ↗
              However, meta is engaged in multiple legal cases about them torrenting copyrighted material. And they bring the reasonable defense that they might have been caught torrenting, but were not specically caught seeding&#x2F;uploading. For all anyone can prove they could have blocked all uploads from their torrent client

              It will be interesting to see if any of those cases reach a judgement on that point, instead of both parties just settling

        2. Trusteando · · focus · HN ↗
          Reading isn&#x27;t the right comparison. Human memory is lossy and fades while LLM encoding is durable with no degradation. The valuable content that the author provides: the content, style, selection of topics, and more is encoded, written into the LLM weights, and they obtain profit from them (now directly, via ads). No one can compete with that kind of copying and pasting from copyright-protected material. And the scale is what hurts authors most: flooding the market with millions of copies on demand, without paying for it.
          1. qarl · · focus · HN ↗
            &gt; No only reading, because the content, style, selection of topics, and more is encoded, written, in the LLM weights

            Exactly analogous to a human reading the material.

            1. asutekku · · focus · HN ↗
              Most people can&#x27;t recite a book they&#x27;ve read verbatim. However, if you ask an LLM to continue a random sentence from a semi-popular book, it can sometimes provide the exact text, unless the system flags the response.
              1. qarl · · focus · HN ↗
                &gt; Most people can&#x27;t recite a book

                Yes - but some people can.

                Are they criminals for reading books?

                1. asutekku · · focus · HN ↗
                  Those some people can&#x27;t recite a book to any given person in the world at any given time. There&#x27;s a difference.
                  1. qarl · · focus · HN ↗
                    So if we invent a way for those people to talk to everyone in the world - then they would be a criminal for reading books whether they did or not?

                    You&#x27;re not making sense.

                    The problem is the reciting. Not the reading. And hence, not the training.

                    1. asutekku · · focus · HN ↗
                      If you would go on a tv and read out loud a book you have not purchased rights to present, according to the most copyright laws in the world, yes you would and you would get a fine. Similarly if you broadcast a tv-show you have not purchased rights would.

                      Whether this is right or not is a seperate question.

                      1. qarl · · focus · HN ↗
                        &gt; If you would go on a tv and read out loud a book

                        That is not the situation we are discussing. No one is arguing that what you describe is infringement.

                        What we are discussing is the training - which happens BEFORE the broadcast. It is analogous to reading. Is simply READING the material an infringement.

                        1. asutekku · · focus · HN ↗
                          Using analogies like &quot;reading&quot; to describe AI training is quite misleading imho. Training effectively embeds the book&#x27;s contents into the model&#x27;s weights. Changing the format doesn&#x27;t change the content; a better analogy is distributing a book&#x27;s text within software.

                          Current copyright laws are simply not prepared for this unprecedented use.

                          1. qarl · · focus · HN ↗
                            &gt; Training effectively embeds the book&#x27;s contents into the model&#x27;s weights. Changing the format doesn&#x27;t change the content

                            Reading a book embeds its contents into your brain. And yet, that is considered fair use.

                            I agree, the analogies are meaningless in a legal context. In the legal context, the courts disagree with you.

                            1. efreak · · focus · HN ↗
                              I&#x27;m not disagreeing with you, but the comparison is breaking down here. Reading a book is the specific intended purpose. It&#x27;s why the book exists, not just fair use. If you couldn&#x27;t remember what was going on in the book as you read it, it would be worthless.
          2. consensus1 · · focus · HN ↗
            Please give me the location of any copyrighted work in the weights of any open source model or a prompt that will retrieve it.
            1. echoangle · · focus · HN ↗
              Go to Deepseek and ask it “Give me the lyrics for Männer by Grönemeyer”.

              It will give you the complete song lyrics which are under copyright.

              1. qarl · · focus · HN ↗
                Yes, and that is obviously copyright infringement.

                The infringement occurs at the time of copy, not at the time of training.

                1. ChickeNES · · focus · HN ↗
                  That begs the question though, should it be? If you’re looking up the lyrics to a song you’ve bought, does it really matter, at least philosophically, how one retrieves them?
                  1. consensus1 · · focus · HN ↗
                    It shouldn&#x27;t be. The idea that writing down the lyrics to a song and putting them on the internet for no financial gain at all should be punishable by some big record label is the kind of over zealous IP protection that 99% of HN users would have been against until something changed a few years ago and this place became a haven for copyright extremists.
                  2. riskable · · focus · HN ↗
                    The law here is no: If you own a copy of the song, you can reproduce the lyrics in whatever way you want.

                    What you can&#x27;t do is distribute the lyrics without the author&#x27;s permission.

                    When you ask DeepSeek to retrieve the lyrics (and it does so), the real question is this: Is DeepSeek merely acting as an intermediary&#x2F;ISP according to the DMCA (in which case they&#x27;d be protected under the safe harbor clauses) or are they illegally redistributing the lyrics without the author&#x27;s permission?

                    Whether or not you own a copy of the lyrics is irrelevant from a legal perspective in this scenario.

                    My guess: If they just retrieved the lyrics from some website and delivered them to you (because you asked), they&#x27;re an ISP. However, if they pulled them out of their own database, they&#x27;re violating copyright.

              2. consensus1 · · focus · HN ↗
                Is this a direct inference call to the model or does it use a web search tool to pull the answer? In the former case I guess that could be infringement. In the latter case it is no more than a tool that facilitates it and would be not be infringement any more than Xerox is infringing if you use their copier to copy the front page of the NYT.
              3. riskable · · focus · HN ↗
                Those lyrics aren&#x27;t in the model. They&#x27;re either in the database DeepSeek makes tool calls into (unlikely), or they&#x27;re out on the Internet and DeepSeek simply retrieved them on your behalf.

                LLMs are far too lossy to be able to store such lyrics in their entirety. In fact, they&#x27;re not even &quot;lossy&quot; since they&#x27;re not even trying to record such information. They&#x27;re just weights for how likely it is that any given word will come after another.

          3. tomjen3 · · focus · HN ↗
            Human memory decays; similarly, LLMs do not have perfect recall. But even if I reread a book every month for 60 years and use the knowledge in there to build a 25 billion dollar empire, the only thing the author is ever going to get from me is the $10.25 the book cost me. And no court in the western world would say he&#x27;s due more.
          4. Trusteando · · focus · HN ↗

            [dead]

        3. zaptheimpaler · · focus · HN ↗
          We have proof that Meta torrented TBs of books off Anna&#x27;s Archive. They didn&#x27;t even buy the books legally in the first place. Fair use doesn&#x27;t magically make piracy legal. The only reason they haven&#x27;t been punished is the pathetic state of our government and court system.
          1. kenmacd · · focus · HN ↗
            Would you be happy if they bought exactly one copy of each book? That might be a dozen sales between the AI labs, some tens of dollars to the author. Is this really what you&#x27;re so upset over?

            And the solution to avoid this piracy has been to buy up rare editions of old books, cut them up and scan them. Better?

            1. zaptheimpaler · · focus · HN ↗
              Excusing illegal behavior when it comes from powerful entities creates a moral hazard and over time leads to the kind of amoral and corrupt government and business leaders we have today. There&#x27;s a pervasive sense in society that money can buy you out of any consequences and morality is basically irrelevant, that&#x27;s how it happens. It&#x27;s the same dynamic as SF failing to police minor shoplifting for years, which led to stores closing and metal bars everywhere that we have today.
              1. kenmacd · · focus · HN ↗
                If this is the issue you actually care about then you&#x27;re in the minority. Most seem to be using it as a proxy for complaining about how the LLM uses the content, not that they didn&#x27;t pay $10 for the content to start with.
                1. zaptheimpaler · · focus · HN ↗
                  People care about rich and powerful entities being proven to be completely immune to any law or consequences, over and over again. It&#x27;s a caste system with money. This issue is widespread and is a centerpiece of midterm elections. Calling it a minority view is a convenient way to disengage.
          2. qarl · · focus · HN ↗
            &gt; Fair use doesn&#x27;t magically make piracy legal.

            Of course it doesn&#x27;t. No one is arguing that.

            We are arguing the bigger issue of whether training is an infringement.

            &gt; The only reason they haven&#x27;t been punished

            The reason they haven&#x27;t been punished is because copyright infringement is not a criminal offense, it&#x27;s civil. And the labs are settling those cases.

        4. throwa356262 · · focus · HN ↗
          Yet &quot;distillation&quot; which is essentially asking someone whose purpose is to answer questions is considered theft.
          1. qarl · · focus · HN ↗
            &gt; is considered theft

            Not by the courts it&#x27;s not.

            I think it might be against the terms of service.

          2. riskable · · focus · HN ↗
            Whoa there: Distillation isn&#x27;t illegal. Who said it was?

            That has yet to be decided in court. At best it&#x27;s a violation of the terms of service, but the only restitution for such a violation is termination of service and by then it&#x27;s too late.

            Remember: Distillation is literally just giving the AI a prompt and seeing what it spits out. If that were illegal, we&#x27;d all be guilty whenever we used AI.

            Examining your competitors doing business is as old as business.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.