‹ BackHN Continuity

Thread

Unsealed Briefs in Authors’ Case v. Microsoft/OpenAI

631 points · 623 comments · papergirl

  1. Skyy93 · · focus · HN ↗
    This is a lobby organisation using only the pieces and bits they like to push their own agenda.

    "sketchy russian website", how about using some more clear description like: A library for sharing books and articles that should be partly public domain because they were paid for by the public. Only some of the material is copyrighted by authors. However, some of their work is so old that it is not reprinted anyway.

    But of course such an explanation would not click.

    I also don't see a problem with statements about making people jobless. Imagine if every robotic or automation company advertised like this: Yeah, you'll buy tons of expensive robots and still rely on expensive labor from real people without any efficiency gains.

    1. lensecat · · focus · HN ↗
      "A lobby organisation?" Of course a single author would not be able to afford facing a multi billion dollar company on their own? And the "sketchy russian website" quote is from OpenAI employees themselves? What are you on about?
      1. Skyy93 · · focus · HN ↗
        First, it's still a lobbying organization, so it's their job to make exaggerated claims, like a union in a company or any other organization with a political purpose. My first point was to highlight that it's not neutral or news related. It's fine that they have their opinion, but it's also my right to say that they're biased.

        The second thing underscores my point. They use one line and think they've made a great point because one employee called LibGen sketchy. This site has been around since the 2010s, and it has helped many people do research. It's not just a sketchy website that suddenly appeared and is always doing bad things. I think a more nuanced stance is necessary.

        1. latexr · · focus · HN ↗
          > They use one line and think they've made a great point because one employee called LibGen sketchy.

          No, you are using one line from the post to discredit them. The release has more than that and it’s not the only communication they made on this matter nor is there any indication it will be the last, it’s just the current one.

          1. Skyy93 · · focus · HN ↗
            They are having it in the subtitle. In general their whole article is about two main points, first the use of stuff from Libgen, second the points of making people jobless. Three of their points are about the jobless thing two about the LibGen.

            About LibGen, there might be more discussion - fair. However, the second argument is no real discussion IMO. Why is putting people out of work suddenly a bad thing? Since when do we argue this when talking about automation?

            1. TeMPOraL · · focus · HN ↗
              There is nothing to discuss about LibGen, really. I don't read this as OpenAI employees even believing LibGen is sketchy. They were worried about optics, because LibGen itself is Russian and does look a bit sketchy, and at the time - much like today - it was easy to make it a headline that makes people pattern-match to "troll farms".

              (And then Russia invaded Ukraine, turning any association with .ru things into potential corporate suicide.)

              Really has nothing to do with LibGen or with OpenAI. It's about people being easy to manipulate into believing bullshit, which is a reasonable worry, and the Authors Guild is trying to do that exact thing OpenAI was worried about.

              1. Skyy93 · · focus · HN ↗
                I agree with you, but I acknowledge that some people (not myself) might feel different about copyright.
                1. TeMPOraL · · focus · HN ↗
                  I acknowledge that, but this part isn't even about copyright!

                  The whole "sketchy russian website" bit resolves entirely about being seen as associated or supporting troll farms and Putin.

        2. bdauvergne · · focus · HN ↗
          You can presume they are biased, then you have to prove it. You cannot assume all they say is biased, no debate is possible with that kind of assumptions.
    2. abroszka33 · · focus · HN ↗
      Any source that this was a library? Even then that would still raise a question if OpenAI is a Russian organisation or not to access that library with good faith.

      I think they just used a Russian torrent site.

      1. Skyy93 · · focus · HN ↗
        They mention it in the text, its this: <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Library_Genesis" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Library_Genesis

        &gt;Microsoft knew about OpenAI’s use of LibGen as early as April 2019

      2. spwa4 · · focus · HN ↗
        Just go look: <a href="https:&#x2F;&#x2F;libgen.im&#x2F;" rel="nofollow">https:&#x2F;&#x2F;libgen.im&#x2F; (may be blocked by your isp, just find an alternate URL somehwere)

        (note: <a href="https:&#x2F;&#x2F;z-library.sk&#x2F;" rel="nofollow">https:&#x2F;&#x2F;z-library.sk&#x2F; is prettier&#x2F;nicer)

      3. [deleted] · · focus · HN ↗

        [deleted]

      4. ZeWaka · · focus · HN ↗
        They&#x27;re talking about LibGen, per the article.
      5. Andrew_nenakhov · · focus · HN ↗
        apparently, the website in question is libgen.io (appears to have been taken down now), currently it has lots of mirrors like libgen.im , libgen.com.de, etc.
    3. 4ndrewl · · focus · HN ↗
      &quot;This is a lobby organisation using only the pieces and bits they like to push their own agenda.&quot;

      Do you have any evidence of them being a lobby organisation (as opposed to OpenAI for example which spends millions of dollars hiring actual lobbyists)

      1. Skyy93 · · focus · HN ↗
        Simply look at the Wikipedia article?

        &gt;The group lobbies at the national and state levels on censorship and tax concerns, and it has initiated or supported several major lawsuits in defense of authors&#x27; copyrights.

      2. TeMPOraL · · focus · HN ↗
        &gt; Do you have any evidence of them being a lobby organisation

        Have you looked at their name?

        1. thatsabadlook · · focus · HN ↗
          It is funny how hard people(?) or maybe bots shill OAI on here. That said why they care at all about some clown like myself &quot;anonymous&quot; opinion on HN is crazy. HN is not an illustrious mind share? A lot of the people here get butterflies thinking about laying off most of their staff for something cheaper. Hell some of the people here get the tinglies if something might be cheaper.

          I&#x27;m more annoyed by the constant whack engagement bait that gets posted here and everywhere about AI. Here I am again engaging still not buying a subscription...

        2. 4ndrewl · · focus · HN ↗
          Never heard of them, but yes they are a lobby company. So they&#x27;re biased and lobbying for their interests and OAI are biased and lobbying for their interests. That adds a lot of colour.
    4. quaintdev · · focus · HN ↗
      It&#x27;s just sad that this is the top comment on Hacker News. Why are we giving free pass to these tech companies? Why are we trusting these CEOs when they have repeatedly broken laws? Remember Aaron Swartz and the fate he suffered? Why is big tech getting away with so much more?
      1. Skyy93 · · focus · HN ↗
        Then you misunderstood my point. I think copyright law should be significantly changed and the current system hurts us all.
        1. ares623 · · focus · HN ↗
          Sure. But the fact is they broke current existing laws, with known punishments with precedents. Same as a new law doesn&#x27;t retroactively punish someone, then a new law shouldn&#x27;t absolve someone before it&#x27;s passed.
          1. Skyy93 · · focus · HN ↗
            &gt;A copyright is a type of intellectual property that gives its owner the exclusive legal right to copy, distribute, adapt, display, and perform a creative work, usually for a limited time.

            The more interesting question is IMO if AI training actually falls into one of these cases. You can read a book and also copy it, but you do not do because of the law. However, you have the ability to do so. Is having the ability to do something already forbidden?

            1. kenmacd · · focus · HN ↗
              In the Anthropic court case it was determined not to, ie that training was transformative.

              It&#x27;s like if you read a plumbing book and then made YouTube videos on how to fix a sink.

          2. ben_w · · focus · HN ↗
            They did. They were found guilty. The case I looked at* was a civil case so this was settled out of court before the court imposed a settlement.

            The law they broke was pirating the materials, not training per se, even though training is what so many people object to: the judge ruled that actually training a model, when the materials you used were ones you otherwise had lawful access to, was not a breach of law.

            IMO, the laws need to change to reflect what tech can now do. This wouldn&#x27;t be the first time, copyright law has had to shift several times before as new means of reproduction are created.

            * the Anthropic one

        2. cisc · · focus · HN ↗
          &gt; the current system hurts us all

          How? The current system enables the GPL. The GPL protects many open source projects.

          1. ChickeNES · · focus · HN ↗
            All the GPL does is restrict rights, unlike permissive licenses.
            1. cisc · · focus · HN ↗
              So.. Linux is bad and we should all use BSDs?

              Why has restrictive Linux succeeded far more than any BSD ever has?

              1. ChickeNES · · focus · HN ↗
                Why yes, I did order a new strawman, thanks so much!
                1. cisc · · focus · HN ↗
                  What strawman? The GPL has been one of the drivers of the success of Linux and copyright is what makes that work. You still haven&#x27;t articulated how that&#x27;s restrictive.
                  1. ChickeNES · · focus · HN ↗
                    &gt;&gt; All the GPL does is restrict rights, unlike permissive licenses. &gt; So.. Linux is bad and we should all use BSDs

                    I did not bring up Linux, much less called it bad, yet you made up a strawman and started attacking it. Good day.

                    1. cisc · · focus · HN ↗
                      We&#x27;re talking about the GPL. Linux is a major project protected by the GPL.

                      A license in isolation isn&#x27;t interesting. The practical results of projects under a license is interesting.

                      Linux is a very successful practical result under the GPL and copyright makes the GPL work. Without copyright the GPL would be unenforceable.

                      1. tpmoney · · focus · HN ↗
                        &gt; Without copyright the GPL would be unenforceable.

                        Without copyright, the GPL would be unnecessary.

                        1. cisc · · focus · HN ↗
                          Without copyright the little guys would constantly be screwed over by the big guys.

                          It really is one of those &quot;without law you cannot have freedom&quot; things.

              2. Macha · · focus · HN ↗
                BSD was under a cloud of a lawsuit at the early stages of corporate adoption of open source operating systems and by the time that was resolved network effects had given Linux too much of a lead.
        3. jrflowers · · focus · HN ↗
          &gt; Then you misunderstood my point.

          No, they were responding to the post defending OpenAI that you wrote. If you meant to communicate something other than “criticism of OpenAI in this context is unwarranted” then it looks like you forgot to do that and wrote something else instead

          1. cwillu · · focus · HN ↗
            Selective and out-of-context quoting isn&#x27;t “criticism of OpenAI”, the same way that saying “checkmate” when your opponent isn&#x27;t in check isn&#x27;t chess.
            1. jrflowers · · focus · HN ↗
              &gt; Selective and out-of-context quoting isn&#x27;t “criticism of OpenAI”, the same way that saying “checkmate” when your opponent isn&#x27;t in check isn&#x27;t chess.

              Is quoting “criticism of OpenAI” without the rest of the post a way of saying “checkmate”?

      2. TeMPOraL · · focus · HN ↗
        &gt; Remember Aaron Swartz and the fate he suffered? Why is big tech getting away with so much more?

        Why are you turning him into perpetuum mobile in his grave?

        Do you really believe Aaron would be arguing against AI companies and for publishing &#x2F; recording guilds on the grounds of intellectual property claims?

        No, it&#x27;s the tech community that did a sudden about-face, and is now all &quot;friendship ended with free access to information and technologies enabling people; now RIAA is my best friend&quot;, and this move is as dumb as that meme (<a href="https:&#x2F;&#x2F;imgflip.com&#x2F;memegenerator&#x2F;137501417&#x2F;Friendship-ended" rel="nofollow">https:&#x2F;&#x2F;imgflip.com&#x2F;memegenerator&#x2F;137501417&#x2F;Friendship-ended).

        1. quaintdev · · focus · HN ↗
          &gt; Do you really believe Aaron would be arguing against AI companies and for publishing &#x2F; recording guilds on the grounds of intellectual property claims?

          That&#x27;s not the point. All the rules and laws are enforced when its you and me but when it&#x27;s big tech the laws are treated by these companies as mere instructions.

          &gt; friendship ended with free access to information and technologies enabling people; now RIAA is my best friend

          Big tech will enable access to free information and will help people reach new heights. Do you see how wrong that sounds?

          1. yorwba · · focus · HN ↗
            The article is about a lawsuit targeting OpenAI. That&#x27;s how rules and laws are typically enforced.
        2. applfanboysbgon · · focus · HN ↗
          Suppose you have three propositions:

          A: &quot;Information is free&quot;

          B: &quot;Information is not free&quot;

          C: &quot;Information is free only for the rich and not free for everyone else, giving the rich a material advantage over everyone else that not only entrenches but accelerates wealth inequality and impedes class mobility&quot;

          You, or Swartz, are an advocate for A. Why, exactly, do you think that obliges you&#x2F;Swartz to prefer C over B while A is not true?

          1. simianwords · · focus · HN ↗
            Sure it’s not free for anyone and both companies and individuals are treated similarly. It’s not like you will be jailed for pirating movies. And neither should OpenAI. What part of this is hard to understand
            1. applfanboysbgon · · focus · HN ↗
              &gt; It’s not like you will be jailed for pirating movies.

              We&#x27;re literally talking in a thread about someone who committed suicide because the US government was hellbent on ruining his life with a felony conviction for piracy.

              1. roenxi · · focus · HN ↗
                Sounds like a gross misapplication of the law enforcement system. The free exchange of information should be legalised.

                As we can see in the good article, the copyright system might almost have cost us a lot of AI capabilities. The damage it has done in cases where the lawyers got ahead of the builders is incalculable.

              2. simianwords · · focus · HN ↗
                His case was materially different

                1. Unauthorised network access

                2. Intent to distribute licensed material

                The labs aren’t doing this. They are doing something similar to you and I downloading torrents. Look, I also think laws should apply somewhat equally to individuals and companies. But this is different.

                1. applfanboysbgon · · focus · HN ↗
                  &gt; Intent to distribute licensed material

                  According to the people trying to ruin his life. This is an accused thought crime, not something he actually did. Also, supposing it actually was something he did, your argument is that building a trillion dollar business on stolen licensed material is legally permissible but giving it away for free is worthy of your life being ruined. Wonderfully coherent world view, that is. Piracy is fine, but only if you hoard it to yourself and profit from it!

                2. shakna · · focus · HN ↗
                  They certainly seem intent on producing and distributing materials based on that licensed material, and stripping out the licensing and pretending it never existed.

                  That doesn&#x27;t seem different to me. Copyright, inherits.

                  If its fair use, like the companies currently claim, maybe it is different. But I suspect that the insane push to create a copyright carveout in countries around the world, is not unrelated to a judgement that it probably isn&#x27;t fair use.

                  1. consensus1 · · focus · HN ↗
                    Give me one example of something distributed by OpenAI that is copyrighted by a third party.
                    1. echoangle · · focus · HN ↗
                      <a href="https:&#x2F;&#x2F;www.hsfkramer.com&#x2F;notes&#x2F;ip&#x2F;2025-11&#x2F;munich-court-finds-copyright-infringement-of-song-lyrics-memorised-by-chatgpt" rel="nofollow">https:&#x2F;&#x2F;www.hsfkramer.com&#x2F;notes&#x2F;ip&#x2F;2025-11&#x2F;munich-court-find...

                      Song lyrics for example

                    2. Eddy_Viscosity2 · · focus · HN ↗
                      Their models, all of them. This has been shown repeatedly when models can return, verbatim, copyrighted material when asked. New York Times is suing over this.

                      It would be like if I sold you a pizza, but when you opened the box it contained all the source code for the latest GTA. The pizza wasn&#x27;t copyrighted by anyone, but the what was inside the box was.

                      1. tpmoney · · focus · HN ↗
                        A xerox machine can produce verbatim copyrighted works when asked as well. That doesn’t make distributing the xerox machine the same as distributing the copyrighted works.

                        Strong IP advocates have argued for years that devices that can be used to infringe copyright are themselves infringement of copyright. So far that hasn’t held up to court analysis provided that device can be and is also used for non-infringing purposes. Given that so far the courts have found that training an AI model is sufficiently transformative to qualify as fair use, it doesn’t seem likely that distributing a model counts as distributing copyrighted material.

                        1. Eddy_Viscosity2 · · focus · HN ↗
                          Your analogy would only be correct if I when bought a xerox machine and brought it home, I could just ask it to print out the copyrighted works without me having them to put on the glass. It&#x27;s not making a copy from one I already had, but providing me a copy when I didn&#x27;t have the original.

                          &gt;training an AI model is sufficiently transformative to qualify as fair use

                          This is the key question and that courts have decided this way so far doesn&#x27;t mean its the correct decision. If the model can encode the copyrighted material with sufficient fidelity to reproduce them on command, it stops being fair use or should anyway.

                          1. tpmoney · · focus · HN ↗
                            &gt; Your analogy would only be correct if I when bought a xerox machine and brought it home, I could just ask it to print out the copyrighted works without me having them to put on the glass. It&#x27;s not making a copy from one I already had, but providing me a copy when I didn&#x27;t have the original.

                            &gt; ...

                            &gt; If the model can encode the copyrighted material with sufficient fidelity to reproduce them on command, it stops being fair use or should anyway.

                            That may be so, IF you could actually do that. Yet in over 200 individual allegations in the Authors Guild vs. Open AI case complaint[1], not a single one of them alleges that you are able to do this. They allege that you could at one point get detailed verbatim quotations, but also note that the models have been explicitly blocked from doing this. Instead, the vast majority of the actual complaints in the case are about generating &quot;summaries&quot; that contain information not in other publicly available summaries, and generating &quot;detailed outlines&quot; of supposed future installments of the copyrighted works, using the characters and details of the story. In other words, all the actual alleged infringements are about either:

                            A) the copies made in order to train the model

                            B) infringing derivative works generated by prompting the AI

                            C) the verbatim copies made from sources to which OpenAI did not have rights to

                            You would think if any of the authors at all had been able to print out verbatim copies of their works without having to take knowing and direct action to circumvent the blocks in place to prevent that from happening, those would have been some of the top complaints in the case. The same held true for Bartz vs. Anthropic, where the judge even noted in his ruling that while generating verbatim copies might indeed be infringement, the plaintiffs never alleged that had happened or was possible.

                            And since Bartz vs. Anthropic has (reasonably IMO) found the training to be sufficiently transformative as to be fair use, the complaints for point A are unlikely to succeed here. The complaints for C almost certainly will succeed, for the same reasons they succeeded against Anthropic.

                            That leaves B. The questions would be:

                            1) Are such &quot;detailed&quot; summaries infringing just because they can include things other summaries have yet to include? Personally, I doubt they&#x27;re going to get much traction here unless the courts split and they win on point A. The fact that other summarizers have left certain details out does not inherently make a new summary with other details an infringing work. If we imagine a world where Empire Strikes Back is a new movie, if none of the public reviews of the movie reveal the twist, but you can ask an AI model to summarize the movie and the AI model reveals the twist, that might be disappointing, but I don&#x27;t think there&#x27;s any argument to be made that it is copyright infringement.

                            2) Are speculative outlines of future unpublished work based on the information in a published work infringing just because they have been created?Is a model that CAN be used by a user to intentionally create an infringing derivative work itself an infringing product? Again without splitting the courts and winning on &quot;training is infringing therefore all outputs are also infringing&quot; I just don&#x27;t see how they can win here. A speculative outline of future works is something people have been doing forever (see also any fan site on the internet). While attempting to publish that outline commercially or produce a new work from that outline might itself be infringement, that infringement is the result of explicit and knowing actions of the user akin to putting a book on a xerox machine and producing a cut and paste fan edit from the work. Again the xerox machine is not itself the infringement, and the individual page copies probably are also not infringement until they are used in a specifically infringing way.

                            3) Does the ability of the model to theoretically produce verbatim copies of the training material if OpenAI were to re-program the model to remove the blocks they have put in place to do that mean the models are themselves infringing. This is perhaps the most &quot;up in the air&quot; question of the 3, but the law generally doesn&#x27;t award damages on the potential for copyright infringement, only on actual acts of infringement. Handbrake and various DVD copying tools do not ship with the keys necessary to defeat the DVD protection schemes, yet they know how to use those keys and accept such keys provided by the users. As far as I know, no court cases have been brought or succeeded against any distributors of DVD ripping software despite the fact that evading the &quot;anti-infringement&quot; blocks in the software is both trivial and exposed to the end user. Given that evading the &quot;anti-infringement&quot; blocks of OpenAI&#x27;s models is neither trivial nor exposed to the end user, I&#x27;m fairly comfortable saying that again without splitting the courts and winning on point A, the authors guild isn&#x27;t likely to win here either.

                            [1]: <a href="https:&#x2F;&#x2F;authorsguild.org&#x2F;app&#x2F;uploads&#x2F;2023&#x2F;12&#x2F;Authors-Guild-OpenAI-Microsoft-Class-Action-Complaint-Dec-2023.pdf" rel="nofollow">https:&#x2F;&#x2F;authorsguild.org&#x2F;app&#x2F;uploads&#x2F;2023&#x2F;12&#x2F;Authors-Guild-O...

                      2. consensus1 · · focus · HN ↗
                        Please provide me an example prompt to do so or some algorithm to extract from the weights for an open source model. I will accept any copyrighted work, any model.
                        1. Eddy_Viscosity2 · · focus · HN ↗
                          You are free to read about exactly this in the NYT lawsuit.
                3. mtlmtlmtlmtl · · focus · HN ↗
                  On what planet is it similar? The labs are taking the copyrighted material and attempting to make billions of dollars from it. And this, in your mind, is &quot;similar&quot; to me downloading a TV show simply to watch it?
                  1. tpmoney · · focus · HN ↗
                    There&#x27;s an argument to be made that you downloading a TV show to watch it is actually worse than what the AI companies are doing. The purpose of making the TV show is to make an entertaining product that people will pay money for in some fashion in order to watch the show. There is no reasonable belief that your act of piracy could ever be a &quot;fair use&quot; of the material. The &quot;social contract&quot; as it were is that if you watch the show, you pay.

                    By comparison, the model isn&#x27;t &quot;watching&quot; the show, as so many people are quick to point out that the &quot;learning&quot; analogy for what AIs are doing is flawed. There was never an intent by the creators that the show would be used to generate mathematical probabilities and weights in a statistical model and no one is deriving entertainment from making the statistical model. I suppose perhaps someone derives entertainment from AI training, but I suspect the number is small enough that &quot;no one&quot; is a reasonable approximation. So using the show to do so at least has an argument towards fair use. Or if the copy used for training was legally purchased, at least in the US it has the actual legal designation as fair use so far.

                    Don&#x27;t get me wrong, I&#x27;m not saying that we should be returning to the days of the RIAA suing teenagers for their college education funds. But it seems pretty obvious that &quot;pirating copyrighted material to explicitly use that material in the way that the creators of the material envisioned selling to you&quot; is similar to, but arguably worse than &quot;using copyrighted material (pirated or not) in a way not envisioned by the creator of that material to create a wholly different product&quot;. In both cases, the livelihood of the creator is possibly being affected, but one of them is a direct 1 for 1 loss of income while the other (again, if not specifically pirated) is an indirect impact.

          2. michaelt · · focus · HN ↗
            Once upon a time, copyright infringement for personal use was barely a crime, while copyright infringement by a for-profit commercial enterprise was a serious matter.

            The idea being (before the rise of online peer-to-peer piracy) to prosecute the people making bootleg VHSes rather than the people buying them.

            With the rise of these AI behemoths, it seems that rule is now inverted: You can download all the pirated ebooks you want, as long as it&#x27;s for large-scale for-profit commercial use.

            1. consensus1 · · focus · HN ↗
              Nothing is inverted. OpenAI was never engaged in any sort of distribution of pirated copies.
            2. masfuerte · · focus · HN ↗
              It wasn&#x27;t a crime at all. It was a civil matter. Over the last forty years it has been criminalized around much of the world under pressure from the USA.
            3. Aunche · · focus · HN ↗
              &gt; You can download all the pirated ebooks you want, as long as it&#x27;s for large-scale for-profit commercial use.

              It&#x27;s literally the opposite. Anthropic paid a $1.5 billion settlement. Litigation against OpenAi is still ongoing. Meanwhile, no one has ever been punished just for consuming pirated media.

              1. idiotsecant · · focus · HN ↗
                Oh boy 1.5 billion, that is almost as much as the cost of the cardboard boxes that their GPUs ship in!! Surely this will make a substantial difference and prevent any future occurrence of this crime.

                It would be shocking and outrageous if this was a case of this being &#x27;the cost of doing business&#x27; for the big guy and a life-ending judgement for the little guy. Luckily our justice system is clearly allowing individual citizens the pleasure and honor of eating cake.

                1. Aunche · · focus · HN ↗
                  1.5 billion is an order of magnitude less than what Anthropic just had bought the books in question. And like I said it&#x27;s infinitely more than what any private has ever paid for doing the same thing.

                  Even if AI companies were literally doing the exact same thing as Aaron Swartz but not getting punished for it, that still doesn&#x27;t make Swartz&#x27;s punishment retroactively their fault. If you have a problem with powerful people being powerful, then I suggest you direct your complaints to the heavens.

              2. cgio · · focus · HN ↗
                And to whom will this settlement be paid? To your random blogger or coder on GitHub?
                1. Aunche · · focus · HN ↗
                  How does this have anything to do with what I said?
        3. thefounder · · focus · HN ↗
          Let’s just pretend that “poor” people who die because of our stupid laws (even now as this AI clown show is playing out) do not matter while the billionaires running these AI companies get a free pass and it’s OK.

          Fun fact: Kim Dotcom is still fighting extradition while these drama queens (I.e Dario) are lecturing us about how much access the peasants should get to AI models fed and trained with stolen IP.

        4. vaylian · · focus · HN ↗
          &gt; Do you really believe Aaron would be arguing against AI companies and for publishing &#x2F; recording guilds on the grounds of intellectual property claims?

          I strongly believe Aaron would oppose the appropriation of content. The problem with AI (in this context) is not that the AI companies gain access to information that regular people can&#x27;t freely access. The problem is that AI erases the information about who originally created a piece of work.

          When people want to freely share their work, then they usually reach for the Creative Commons licenses and not for Public Domain, because the latter doesn&#x27;t protect authorship.

          1. smugglerFlynn · · focus · HN ↗
            Exactly. Imagine alternative time line where tech is the same but each creative provides their own trained model that scales their own creative vision and skills, and keeps them in control of their own work.

            Imagine OpenAI, Anthropic &amp;co having to compete by hiring [thousands] of their own talent to help training their commercial models.

            Another thing is attribution. Even a book that was re-published illegally can be easily attributed to its author. What happened is the exact opposite: no attribution, not even a notice, just obfuscation that strips away any traces of the original ideas and original work. Imagine piracy websites and trackers just dropping first few pages that name their authors, and publishing &quot;the book you are looking for&quot;. This is exactly what happened.

            1. doginasuit · · focus · HN ↗
              &gt; each creative provides their own trained model that scales their own creative vision and skills

              I like the vision, but the model would need more than just the author&#x27;s work. The situation you are imagining would require models like we have as a base.

            2. tpmoney · · focus · HN ↗
              &gt; Imagine alternative time line where tech is the same but each creative provides their own trained model that scales their own creative vision and skills, and keeps them in control of their own work.

              This is a timeline that can not and will never exist if the Authors Guild and most of the other anti-AI lawsuits succeed. Because if they do succeed, the only people who will be able to provide an AI model will be companies with enough resources to license all the training data in perpetuity. The Authors Guild isn&#x27;t mad because the authors can&#x27;t train and use their own AIs, they&#x27;re mad because the AI companies are making money and they&#x27;re not getting what they perceive as their fair cut of that money.

              &gt; Imagine OpenAI, Anthropic &amp;co having to compete by hiring [thousands] of their own talent to help training their commercial models.

              If that&#x27;s what the AI companies would need to do, how then does &quot;each creative&quot; compete? What authors or artists do you know that can afford to hire &quot;thousands&quot; of people to help them build bespoke AI models?

              It seems to me that we should be figuring out how to make public models and datasets that can be used by anyone, not further strengthening copyright so that AI models can only be produced by companies with the resources to hire thousands of people.

        5. yqx · · focus · HN ↗
          Aaron Swartz fought to &quot;free&quot; knowledge and information produced by (in many cases publicly funded) research that was and still is hoarded by the academic publishing industry so that they can profit of it.

          What big AI companies have done is illegally hoovering up copyrighted creative output of individuals and creating a situation where the wages that normally would be paid to those individuals instead go to that one company (that stole their work) which now becomes disproportionally rich and powerful.

          In both cases companies obtain money and power by hoarding information obtained through dubious means (in the former case most academics willingly participate while at the same tone they often don&#x27;t really have a choice). Exactly what Swartz was fighting against.

          1. bezier-curve · · focus · HN ↗

            [dead]

            1. Macha · · focus · HN ↗
              Obviously it&#x27;s hard to predict the opinions of the dead, and I think a lot of the disagreement here comes from whether they view his ideas as being motivated by anti-copyright as a standalone idea or part of a larger pro-little-guy stance. I think depending on which you think it was, you can come to a different idea of what his view would be on open models. But surely both views would lead to an anti-closed models stance? Would a hypothetical generation later Aaron want to publish the weights for Claude rather than academic publisher articles?
              1. bezier-curve · · focus · HN ↗
                I am not talking about open models. I am talking about proprietary models being trained on public information. A lot of people want to posture about what an ideological person might think, and some of it seems conveniently positioned to use a past figure that can&#x27;t speak for themselves.

                I think copyright has real intrinsic value in our society, but having been subjected to threatening legal letters for legit fair use situations in my youth, as a result I don&#x27;t think the legal structure behind it is sane. I always respect copyright when I can, but it&#x27;s gotten to a point where fair use is not equally considered.

              2. parineum · · focus · HN ↗
                &gt; Obviously it&#x27;s hard to predict the opinions of the dead

                So we shouldn&#x27;t.

      3. MattGaiser · · focus · HN ↗
        We tend to accept lawbreaking if it is for a product we want.

        There are several multi billion dollar companies where the founding thesis was “what if we just ignore the law?”

      4. simianwords · · focus · HN ↗
        Aaron was sued for distributing and not for pirating. They are different offences
        1. RossBencina · · focus · HN ↗
          As far as I know Aaron never distributed that which was collected in the MIT data closet.
      5. cwillu · · focus · HN ↗
        Calling out equivocation and out-of-context quotes is not “giving a free pass”. If there is a case to be made (and to be clear, I believe there is), then shouldn&#x27;t be resorting to rhetorical slight-of-hand to “prove” it.
      6. sevenzero · · focus · HN ↗

        [dead]

      7. nl · · focus · HN ↗
        &gt; Aaron Swartz and the fate he suffered?

        I think what Swartz did was moral, and his prosecution was unjust.

        I think what OpenAI did was moral, and them getting sued for it is unjust.

        Why do you have one position for Swartz and a different one for OpenAI?

        (Aaron Swartz was a mailing-list friend of mine, so I do have some bias here. But in part we knew each other because our moral position on this was similar)

        1. SXX · · focus · HN ↗
          &gt; Why do you have one position for Swartz and a different one for OpenAI?

          I think many people on HN dont mind OpenAI use of copyrighted material, but do not like how they try to do regulatory capture of a market and try to say their own copyright is now somehow more important.

          Like how OpenAI trying to make &quot;distillation&quot; illegal while it exactly what they did with whole intetnet, books, everything.

          1. nl · · focus · HN ↗
            I agree entirely with this position!

            But it is actually possible to agree with one thing a company does and disagree with others.

        2. michaelt · · focus · HN ↗
          Imagine a town where (a) a corrupt cop issues tickets for fake traffic violations; but (b) he doesn&#x27;t do that to his fellow cops, or the mayor&#x27;s friends.

          Does (b) make things more just, as certain possible unjust acts don&#x27;t happen?

          Or does it compound the injustice by creating a double-standard, and perpetuate it by hiding the problem from any with the power to bring an end to it?

        3. kannanvijayan · · focus · HN ↗
          There seem to be two groups of people talking past each other here, which is worth acknowledging.

          On the one hand, there is a notion of consistent principles regarding the legal handling of the topic, which is being appealed to by some people such as yourself.

          The other side seems to be drawing attention to the fact that these principles are not applied consistently by society. They&#x27;re questioning the moral validity of holding to principle in a circumstance where it&#x27;s guaranteed to be applied with very specific biases that are rarely explicitly stated.

          I&#x27;ve noticed this talking-past happening in other subjects too. For example American drug laws. There&#x27;s the principle of drug laws, and then there&#x27;s the practice of which kinds of people gets the laws applied to them. One group of people focus on the one, and another focus on the other, and they just kind of talk past each other.

          1. nl · · focus · HN ↗
            I think this is a fair point, except that in this case the same laws[1] are being applied to both, and the person I was replying to was literally arguing both ways!

            [1] broadly the same laws: I do understand one was criminal and one is civil and yes I agree that injustice. To me neither should be criminal.

        4. [deleted] · · focus · HN ↗

          [deleted]

        5. isityettime · · focus · HN ↗
          &gt; Why do you have one position for Swartz and a different one for OpenAI?

          &quot;Why do you have one position for an activist and another for a eight-hundred and fifty-two billion dollar, for-profit corporation?&quot;

          &quot;Why to you have one position for someone who wanted to grow the intellectual commons and another for a corporation trying to enclose it?&quot;

          &quot;Why do you have one position for someone who gave his work away for free and another for a company that charges for access to proprietary tech?&quot;

          1. nl · · focus · HN ↗
            I think it&#x27;s good when eight-hundred and fifty-two billion dollar, for-profit corporation establish legal precedent than can then be used to protect activists.
            1. ChickeNES · · focus · HN ↗
              And archivists!
              1. isityettime · · focus · HN ↗
                Isn&#x27;t the whole basis of the legal success of the AI companies so far that training an LLM is &quot;transformative&quot;, i.e., something that would not apply in any way to archival?
        6. frtt2 · · focus · HN ↗
          Your posts lack obvious nuance.

          What’s the motive? For openAI its profit.

          1. nl · · focus · HN ↗
            I agree that OpenAI&#x27;s motive is profit.

            I stand by my moral position.

      8. blini-kot · · focus · HN ↗
        why?

        uhhh... capitalism

      9. mannanj · · focus · HN ↗
        We aren&#x27;t. Most people are silent majority, just not sharing their true opinions out of fear, laziness, or some other reason like maybe just the bias of how most living entities function and also related to how the bystander effect also works.

        Then, who is &quot;we&quot; here giving the appearance of a consensus opinion here and in mainstream? It&#x27;s a vocal minority, it&#x27;s the powerful, it&#x27;s the causes they fund and put resources behind to continue to preserve their causes. And now today, it&#x27;s astroturfing, fake AI-LLM-bots almost indistinguishable from you and I. Don&#x27;t mistake artificial consensus for reality.

        &gt; Why is big tech getting away with so much more?

        IMO Because the majority of people are passive, standing by, tolerating abuse and trickery by the minority. This is a perpetual cycle in humanity: those minority use their power and leverage and abuse their positions until they are ousted. We have tolerated this because we haven&#x27;t stood up yet and acted to change things and demand equal enforcement of the laws that appear to apply to us but not to them. If history says anything, they are afraid and panicking and will continue to be more abusive until they push our buttons more and more, and usually it explodes in their face because they still need us (which is why mainstream rich people push robotics and automation and AI down our throats so aggressively because they know all this) and yet never have minority humans won that approach before. Leadership always changes. Life always changes, and no force can stay dominant for ever.

      10. riskable · · focus · HN ↗
        Because it&#x27;s everyone&#x27;s right to scrape the Internet and save what you find?
    5. [deleted] · · focus · HN ↗

      [deleted]

    6. trompetenaccoun · · focus · HN ↗
      It can be a bit confusing due to the terrible style of the article (ironic given the source) but it seems the &quot;sketchy russian website&quot; part is a direct quote by Anthropic&#x27;s Sam McCandlish. And apparently Dario Amodei referred to it as sketchy as well.

      I find the brazenness of saying this while running what&#x27;s arguably the largest copyright theft operation in human history astonishing. If libgen is &quot;sketchy&quot;, then what is OpenAI?

      1. qarl · · focus · HN ↗
        &gt; the largest copyright theft operation in human history

        Many people think that it was fair use: training is akin to reading, not copying.

        Especially the courts.

        1. trompetenaccoun · · focus · HN ↗
          That&#x27;s not been legally established, the litigation is ongoing. And if mere downloading and reading of copyrighted material were legal, how come torrent users have been fined for it in the thousands?

          The law is the law, there can&#x27;t be different law for corporations with billions in backing. I don&#x27;t agree with current copyright laws btw and think they should be changed. However, they probably should have lobbied for that before illegally downloading all this material.

          1. qarl · · focus · HN ↗
            &gt; That&#x27;s not been legally established

            100% of the rulings agree with me.

            The piracy is not in question. It is unarguably copyright violation.

            But that&#x27;s not what anyone means in this context. Training is what everyone means.

            &gt; The law is the law, there can&#x27;t be different law for corporations with billions in backing.

            I didn&#x27;t say otherwise. That&#x27;s a straw man.

            1. trompetenaccoun · · focus · HN ↗
              So we agree they have violated copyright at a much larger scale than LibGen, yet they call LibGen &quot;sketchy&quot; for doing the same thing? Absurd, what exactly are we arguing here?
              1. qarl · · focus · HN ↗
                &gt; So we agree they have violated copyright at a much larger scale than LibGen

                No.

            2. latexr · · focus · HN ↗
              So, which is it? First you said it was fair use, now you’re saying it’s unarguably copyright violation. You can’t have it both ways.
              1. qarl · · focus · HN ↗
                Training is fair use - the torrenting was infringement. Sorry if that was confusing.
                1. mrdependable · · focus · HN ↗
                  Training has not, in fact, been decided on yet. Of the four criteria deciding whether something is fair use, training only maybe passes two.

                  <a href="https:&#x2F;&#x2F;www.copyright.gov&#x2F;ai&#x2F;Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf" rel="nofollow">https:&#x2F;&#x2F;www.copyright.gov&#x2F;ai&#x2F;Copyright-and-Artificial-Intell...

                  1. qarl · · focus · HN ↗

                    [dead]

                  2. tpmoney · · focus · HN ↗
                    Training in the US has, in fact, been decided (at least to the extent that anything has currently been decided). Bartz vs Anthropic specifically ruled training an AI model on legally owned copyrighted material is sufficiently transformative[1]:

                        This order grants summary judgment for Anthropic that the training use was a fair use.
                        And, it grants that the print-to-digital format change was a fair use for a different reason. But it
                        denies summary judgment for Anthropic that the pirated library copies must be treated as
                        training copies.
                    
                    The document you linked was written a month before the Bartz decision was reached. It&#x27;s also worth noting even the document you linked says this in its conclusion:

                        Various uses of copyrighted works in AI training are likely to be transformative. The
                        extent to which they are fair, however, will depend on what works were used, from what
                        source, for what purpose, and with what controls on the outputs—all of which can affect the
                        market.
                    
                    
                    [1]: <a href="https:&#x2F;&#x2F;copyrightalliance.org&#x2F;wp-content&#x2F;uploads&#x2F;2025&#x2F;06&#x2F;Bartz-v.-Anthropic-Order.pdf" rel="nofollow">https:&#x2F;&#x2F;copyrightalliance.org&#x2F;wp-content&#x2F;uploads&#x2F;2025&#x2F;06&#x2F;Bar...
                    1. mrdependable · · focus · HN ↗
                      That is not binding, hence the NYT trial.

                      That is not what the text you quoted is saying. It says that the output may be transformative, which is one criteria, but the other criteria depends on how it is used and what the source is.

          2. echoangle · · focus · HN ↗
            &gt; And if mere downloading and reading of copyrighted material were legal, how come torrent users have been fined for it in the thousands?

            Because torrenting includes redistribution, since you also upload to other peers

            1. wongarsu · · focus · HN ↗
              However, meta is engaged in multiple legal cases about them torrenting copyrighted material. And they bring the reasonable defense that they might have been caught torrenting, but were not specically caught seeding&#x2F;uploading. For all anyone can prove they could have blocked all uploads from their torrent client

              It will be interesting to see if any of those cases reach a judgement on that point, instead of both parties just settling

        2. Trusteando · · focus · HN ↗
          Reading isn&#x27;t the right comparison. Human memory is lossy and fades while LLM encoding is durable with no degradation. The valuable content that the author provides: the content, style, selection of topics, and more is encoded, written into the LLM weights, and they obtain profit from them (now directly, via ads). No one can compete with that kind of copying and pasting from copyright-protected material. And the scale is what hurts authors most: flooding the market with millions of copies on demand, without paying for it.
          1. qarl · · focus · HN ↗
            &gt; No only reading, because the content, style, selection of topics, and more is encoded, written, in the LLM weights

            Exactly analogous to a human reading the material.

            1. asutekku · · focus · HN ↗
              Most people can&#x27;t recite a book they&#x27;ve read verbatim. However, if you ask an LLM to continue a random sentence from a semi-popular book, it can sometimes provide the exact text, unless the system flags the response.
              1. qarl · · focus · HN ↗
                &gt; Most people can&#x27;t recite a book

                Yes - but some people can.

                Are they criminals for reading books?

                1. asutekku · · focus · HN ↗
                  Those some people can&#x27;t recite a book to any given person in the world at any given time. There&#x27;s a difference.
                  1. qarl · · focus · HN ↗
                    So if we invent a way for those people to talk to everyone in the world - then they would be a criminal for reading books whether they did or not?

                    You&#x27;re not making sense.

                    The problem is the reciting. Not the reading. And hence, not the training.

                    1. asutekku · · focus · HN ↗
                      If you would go on a tv and read out loud a book you have not purchased rights to present, according to the most copyright laws in the world, yes you would and you would get a fine. Similarly if you broadcast a tv-show you have not purchased rights would.

                      Whether this is right or not is a seperate question.

                      1. qarl · · focus · HN ↗
                        &gt; If you would go on a tv and read out loud a book

                        That is not the situation we are discussing. No one is arguing that what you describe is infringement.

                        What we are discussing is the training - which happens BEFORE the broadcast. It is analogous to reading. Is simply READING the material an infringement.

                        1. asutekku · · focus · HN ↗
                          Using analogies like &quot;reading&quot; to describe AI training is quite misleading imho. Training effectively embeds the book&#x27;s contents into the model&#x27;s weights. Changing the format doesn&#x27;t change the content; a better analogy is distributing a book&#x27;s text within software.

                          Current copyright laws are simply not prepared for this unprecedented use.

                          1. qarl · · focus · HN ↗
                            &gt; Training effectively embeds the book&#x27;s contents into the model&#x27;s weights. Changing the format doesn&#x27;t change the content

                            Reading a book embeds its contents into your brain. And yet, that is considered fair use.

                            I agree, the analogies are meaningless in a legal context. In the legal context, the courts disagree with you.

                            1. efreak · · focus · HN ↗
                              I&#x27;m not disagreeing with you, but the comparison is breaking down here. Reading a book is the specific intended purpose. It&#x27;s why the book exists, not just fair use. If you couldn&#x27;t remember what was going on in the book as you read it, it would be worthless.
          2. consensus1 · · focus · HN ↗
            Please give me the location of any copyrighted work in the weights of any open source model or a prompt that will retrieve it.
            1. echoangle · · focus · HN ↗
              Go to Deepseek and ask it “Give me the lyrics for Männer by Grönemeyer”.

              It will give you the complete song lyrics which are under copyright.

              1. qarl · · focus · HN ↗
                Yes, and that is obviously copyright infringement.

                The infringement occurs at the time of copy, not at the time of training.

                1. ChickeNES · · focus · HN ↗
                  That begs the question though, should it be? If you’re looking up the lyrics to a song you’ve bought, does it really matter, at least philosophically, how one retrieves them?
                  1. consensus1 · · focus · HN ↗
                    It shouldn&#x27;t be. The idea that writing down the lyrics to a song and putting them on the internet for no financial gain at all should be punishable by some big record label is the kind of over zealous IP protection that 99% of HN users would have been against until something changed a few years ago and this place became a haven for copyright extremists.
                  2. riskable · · focus · HN ↗
                    The law here is no: If you own a copy of the song, you can reproduce the lyrics in whatever way you want.

                    What you can&#x27;t do is distribute the lyrics without the author&#x27;s permission.

                    When you ask DeepSeek to retrieve the lyrics (and it does so), the real question is this: Is DeepSeek merely acting as an intermediary&#x2F;ISP according to the DMCA (in which case they&#x27;d be protected under the safe harbor clauses) or are they illegally redistributing the lyrics without the author&#x27;s permission?

                    Whether or not you own a copy of the lyrics is irrelevant from a legal perspective in this scenario.

                    My guess: If they just retrieved the lyrics from some website and delivered them to you (because you asked), they&#x27;re an ISP. However, if they pulled them out of their own database, they&#x27;re violating copyright.

              2. consensus1 · · focus · HN ↗
                Is this a direct inference call to the model or does it use a web search tool to pull the answer? In the former case I guess that could be infringement. In the latter case it is no more than a tool that facilitates it and would be not be infringement any more than Xerox is infringing if you use their copier to copy the front page of the NYT.
              3. riskable · · focus · HN ↗
                Those lyrics aren&#x27;t in the model. They&#x27;re either in the database DeepSeek makes tool calls into (unlikely), or they&#x27;re out on the Internet and DeepSeek simply retrieved them on your behalf.

                LLMs are far too lossy to be able to store such lyrics in their entirety. In fact, they&#x27;re not even &quot;lossy&quot; since they&#x27;re not even trying to record such information. They&#x27;re just weights for how likely it is that any given word will come after another.

          3. tomjen3 · · focus · HN ↗
            Human memory decays; similarly, LLMs do not have perfect recall. But even if I reread a book every month for 60 years and use the knowledge in there to build a 25 billion dollar empire, the only thing the author is ever going to get from me is the $10.25 the book cost me. And no court in the western world would say he&#x27;s due more.
          4. Trusteando · · focus · HN ↗

            [dead]

        3. zaptheimpaler · · focus · HN ↗
          We have proof that Meta torrented TBs of books off Anna&#x27;s Archive. They didn&#x27;t even buy the books legally in the first place. Fair use doesn&#x27;t magically make piracy legal. The only reason they haven&#x27;t been punished is the pathetic state of our government and court system.
          1. kenmacd · · focus · HN ↗
            Would you be happy if they bought exactly one copy of each book? That might be a dozen sales between the AI labs, some tens of dollars to the author. Is this really what you&#x27;re so upset over?

            And the solution to avoid this piracy has been to buy up rare editions of old books, cut them up and scan them. Better?

            1. zaptheimpaler · · focus · HN ↗
              Excusing illegal behavior when it comes from powerful entities creates a moral hazard and over time leads to the kind of amoral and corrupt government and business leaders we have today. There&#x27;s a pervasive sense in society that money can buy you out of any consequences and morality is basically irrelevant, that&#x27;s how it happens. It&#x27;s the same dynamic as SF failing to police minor shoplifting for years, which led to stores closing and metal bars everywhere that we have today.
              1. kenmacd · · focus · HN ↗
                If this is the issue you actually care about then you&#x27;re in the minority. Most seem to be using it as a proxy for complaining about how the LLM uses the content, not that they didn&#x27;t pay $10 for the content to start with.
                1. zaptheimpaler · · focus · HN ↗
                  People care about rich and powerful entities being proven to be completely immune to any law or consequences, over and over again. It&#x27;s a caste system with money. This issue is widespread and is a centerpiece of midterm elections. Calling it a minority view is a convenient way to disengage.
          2. qarl · · focus · HN ↗
            &gt; Fair use doesn&#x27;t magically make piracy legal.

            Of course it doesn&#x27;t. No one is arguing that.

            We are arguing the bigger issue of whether training is an infringement.

            &gt; The only reason they haven&#x27;t been punished

            The reason they haven&#x27;t been punished is because copyright infringement is not a criminal offense, it&#x27;s civil. And the labs are settling those cases.

        4. throwa356262 · · focus · HN ↗
          Yet &quot;distillation&quot; which is essentially asking someone whose purpose is to answer questions is considered theft.
          1. qarl · · focus · HN ↗
            &gt; is considered theft

            Not by the courts it&#x27;s not.

            I think it might be against the terms of service.

          2. riskable · · focus · HN ↗
            Whoa there: Distillation isn&#x27;t illegal. Who said it was?

            That has yet to be decided in court. At best it&#x27;s a violation of the terms of service, but the only restitution for such a violation is termination of service and by then it&#x27;s too late.

            Remember: Distillation is literally just giving the AI a prompt and seeing what it spits out. If that were illegal, we&#x27;d all be guilty whenever we used AI.

            Examining your competitors doing business is as old as business.

      2. TeMPOraL · · focus · HN ↗
        They didn&#x27;t believe it was sketchy. They were just worried that the commentariat on HN will frame it in a dumb, manipulative way like that.

        Judging by how AI threads look like for the past year, they were absolutely right to be worried.

        &gt; largest copyright theft operation in human history

        In fact, you&#x27;re doing exactly that right here.

        1. discreteevent · · focus · HN ↗
          &gt; frame it in a dumb, manipulative way like that &gt; In fact, you&#x27;re doing exactly that right here.

          You&#x27;re dead fucking right they are doing exactly that. They are saying that what is wrong is wrong.

          Instead you seem to be making out that AI companies are some kind of victim that has to &quot;worry&quot; about &quot;manipulation&quot;. Meanwhile authors are out of a job right now and not by accident. What&#x27;s up with that?

        2. probably_wrong · · focus · HN ↗
          The comment you&#x27;re replying to is citing almost verbatim [1] Microsoft’s director of Applied Science, Brent Hecht, who called OpenAI&#x27;s data collection practices &quot;the largest theft of labor in human history&quot; in an internal memo.

          [1] <a href="https:&#x2F;&#x2F;techcrunch.com&#x2F;2026&#x2F;09&#x2F;17&#x2F;microsoft-exec-called-ai-scraping-the-largest-theft-of-labor-in-human-history-new-unredacted-filings-reveal&#x2F;" rel="nofollow">https:&#x2F;&#x2F;techcrunch.com&#x2F;2026&#x2F;09&#x2F;17&#x2F;microsoft-exec-called-ai-s...

          1. qarl · · focus · HN ↗
            Yes, one man at Microsoft agrees with you.
          2. TeMPOraL · · focus · HN ↗
            It&#x27;s not citing, it&#x27;s regurgitating, like a stochastic parrot.
            1. throworangeaway · · focus · HN ↗
              It is citing directly from unsealed court document. For an Opinion Haver training here all day, whose only other contribution to the world was animated Nyan Cat in Emacs status bar, you have poor comprehension skills.
        3. latexr · · focus · HN ↗
          &gt; They were just worried that the commentariat on HN will frame it in a dumb, manipulative way like that.

          You’re chastising others for a tone you are yourself employing, and are making monumental assumptions based on a few choice quotes. From the quotes alone you can’t tell if OpenAI thought libgen was sketchy or not.

          Also, contrary to what you’re claiming, they were wrong. HN in general seems to approve of libgen when used for its purpose of downloading some books on an individual level. The complaint you’re replying to is about what OpenAI did with the data, it has nothing to do with the website they got it from.

        4. zaptheimpaler · · focus · HN ↗
          There&#x27;s no actual argument here, just name-calling. Meta literally pirated 80TB of books off of Anna&#x27;s Archive. OpenAI pirated off LibGen as well. You could be prosecuted for doing the same thing, because its illegal. Piracy is illegal, piracy at the scale they committed is extremely illegal. This shouldn&#x27;t be a hard or controversial concept.
          1. ChickeNES · · focus · HN ↗
            Who cares. IP is a fiction, all copyrights and patents should be abolished.
            1. zaptheimpaler · · focus · HN ↗
              Yeah maybe make a few author friends and tell them that they don&#x27;t deserve any payment for their work. I&#x27;m sure the world will be much better once we even further reduce the incentives to do hard intellectual work and increase the incentives to churn out AI slop tiktok videos instead.
              1. ChickeNES · · focus · HN ↗
                If a machine can do the same work for nearly free, &quot;but the existing people want to keep getting paid&quot; is not an argument for copyright.
                1. techpression · · focus · HN ↗
                  It can’t, it needs hundreds of thousands (if not multiple millions) of human hours spent making content for it.
                  1. ChickeNES · · focus · HN ↗
                    Training data? Luckily we have all of human history and culture to digitize to solve that problem.
                2. throwaway173738 · · focus · HN ↗
                  Which is fine if you like human culture to never change for the next ten thousand years. LLMs can only ever interpolate or extrapolate things they’ve already been trained on.
                  1. ChickeNES · · focus · HN ↗
                    I don’t accept the premise, it’s hard to take anyone seriously who is parroting the “stochastic parrot” argument in 02026.
                    1. throwaway173738 · · focus · HN ↗
                      Your assertion that we’ll somehow get brand new human culture from something that isn’t human might be a little shaky too.
    7. latexr · · focus · HN ↗
      “Sketchy Russian website” is part of a quote by Sam McCandlish (who worked at OpenAI), not the Author’s Guild characterisation.

      Also, defending it on the basis that some books on libgen are public domain is a poor excuse, like claiming people use The Pirate Bay to download Linux ISOs. Even if some of that is true, we all know that use case is not the popular one.

    8. probably_wrong · · focus · HN ↗
      &gt; I also don&#x27;t see a problem with statements about making people jobless

      I do see a problem with a company loudly announcing that they are going to make people&#x27;s lives miserable purely for profit. Leaving aside that it goes against OpenAI&#x27;s stated mission (&quot;to ensure that artificial general intelligence benefits all of humanity&quot;), the disdain for the lives they are intentionally trying to ruin makes it a problem.

      And even if you believe that the transition is inevitable, as it is the case with phasing out combustion engines in cars, anyone reasonable would see that the transition is gradual to give people time to adapt. Instead of doing that, OpenAI is burning cash at astonishing rates, polluting the environment, and killing personal computing with the only aim of being the only ones left atop the ruins. I do see a problem with that.

    9. fzeroracer · · focus · HN ↗
      And OpenAI isn&#x27;t a lobby organisation?

      Plus you somehow didn&#x27;t even read the article properly, given that &quot;sketchy russian website&quot; is part of a direct quote.

      1. Skyy93 · · focus · HN ↗
        I have read the article. I called out that they are putting so much emphasis on this term, which is one line, one opinion and build 1&#x2F;2 of their argument: AI-Company - bad! on it.
    10. zaptheimpaler · · focus · HN ↗
      &gt; A library for sharing books and articles that should be partly public domain because they were paid for by the public

      This is some incredible mental gymnastics here, wow. Some books should be public domain (even if they actually aren&#x27;t), and this magically justifies stealing from an 80TB library of a large fraction of every book ever published including millions of books that have no public funding.

    11. lukewarm707 · · focus · HN ↗
      the issue is a matter of principles and hypocrisy.

      you should not enclose the commons.

      that is what openai and anthropic have done; capture the commons, lock the model, restrict the outputs.

      1. tpmoney · · focus · HN ↗
        They have perhaps attempted to lock the model, but so far I don&#x27;t think we&#x27;ve seen any rulings that distillation or using the models to train other models is any less fair use than the original model training the companies have engaged in. Of course, the Authors Guild and other lawsuits against the model makers are arguing that it shouldn&#x27;t be fair use. But if the AI companies prevail, there&#x27;s no real argument to be made that training on their outputs isn&#x27;t itself also fair use.

        Further, I&#x27;m not sure how they&#x27;ve &quot;captured the commons&quot;. By definition the &quot;commons&quot; belongs to us all. Nothing prevents someone else from doing the same thing. That is, unless the Authors Guild succeeds in splitting the courts over the fair use of AI training and the resolution of that split finds that training isn&#x27;t fair use. Then the commons can only be used by companies or people with pockets deep enough to license the material in perpetuity.

    12. bdauvergne · · focus · HN ↗
      If it&#x27;s mainly public domain, AI weights should be too. No more copyright for them they don&#x27;t deserve it.
    13. cornholio · · focus · HN ↗
      Robot companies don&#x27;t appropriate the work of other people. That&#x27;s the fundamental point of contention here: that LLMs cannot be trained without the human labor of the authors; yet, they directly compete with them in the marketplace. I haven&#x27;t yet seen an LLM trained only on public domain material, but its capabilities are likely to be very limited.

      That&#x27;s also the key political compromise underlying the notion of copyright: that someone is entitled to the fruits of their labor, and should not be economically hindered by a product that could not have existed without said work. That&#x27;s the basis on which the &quot;derivative work&quot; copyright doctrine emerged: a work sufficiently original that it does not displace the work on which it is based. LLMs fail to abide by that political compromise by a country mile.

      1. parineum · · focus · HN ↗
        Didn&#x27;t I just read an article about Tesla employees not wanting to train their robot replacements?
      2. tpmoney · · focus · HN ↗
        &gt; Robot companies don&#x27;t appropriate the work of other people.

        I think there&#x27;s a pretty good argument to be made that every single robot is built on the labor, creative and technical knowledge and advancements of the workers that robot replaced. Robots after all, much more than LLMs, are incapable of creative output. Some human (probably a laborer) figured out how to stamp the steel in just the right ways, or how to cut the patterns in just the right ways so that the product could be manufactured. Then a robot company came in and stole that creative output, or more likely was sold that creative output by the company owners who stole&#x2F;bought (depending on your point of view about labor and the ownership of labor&#x27;s creative outputs in the current US legal system) to produce a robot that then displaced the laborer who created the process in the first place.

    14. mrqwen · · focus · HN ↗

      [dead]

    15. contubernio · · focus · HN ↗
      As a long time user (most professional mathematicians have been) of the now mostly defunct Library Genesis, it is certainly more accurate to describe it as a &quot;sketchy Russian website&quot; than as &quot;a library for sharing books and articles that should be partly public domain because they were paid for by the public&quot;. In my country we use it precisely because public universities (paid for by the public) cannot afford to buy math and physics books.
    16. jordemort · · focus · HN ↗
      save some boot for the rest of us
      1. Skyy93 · · focus · HN ↗
        This comment left me chuckling. Thanks for that.
      2. ChickeNES · · focus · HN ↗
        You can’t address the argument, so you resort to ad hominems?
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.