‹ BackHN Continuity

Thread

The AI Race Just Got Awkward

412 points · 463 comments · allisdust

  1. cmiles8 · · focus · HN ↗
    Why is it a problem that the Chinese labs are just distilling down Anthropic’s models? Aren’t Anthropic’s models not just distilling down other people’s work?

    Feels like Anthropic crying do as I say not as I do.

    1. jedberg · · focus · HN ↗
      What Anthropic is doing requires way more resources than what the Chinese labs are doing. So their complaint is that they do 95% of the work and the Chinese labs do the last 5% and call it their own.

      An argument can be made that Anthropic is also only doing the last 5% of the work (because the content they are training on was the other 95%) but that's a bit more philosophical.

      1. JackFr · · focus · HN ↗
        But the analogy still holds.

        The original authors of all the text, creators of the media and developers of the software did far more work than Anthropic.

        1. layer8 · · focus · HN ↗
          SOTA models cost hundreds of millions to train. Did creating the contents of the text corpus they were trained on really cost an equivalent of 20x as much (~10 billions)? I honestly don’t know, but I could imagine it having been significantly less.

          This isn’t meant as a moral argument, just musing about the relative cost comparison.

          1. Ar-Curunir · · focus · HN ↗
            Yes, duh! Human output across the millennia is worth much more than whatever is being invested in frontier labs.

            How is this even a question.

            1. bmacho · · focus · HN ↗
              They are solving unsolved math problems right now, so probably soon or very soon their output will be more valuable than all human recorded knowledge.
              1. mekael · · focus · HN ↗
                Pre vaccination smallpox killed hundreds of millions of people just in the twentieth century [0], the knowledge that allowed for the creation of just that vaccine is worth hundreds trillions of dollars in humans lives, let alone all of `the knowledge and experiences those people were involved in.

                The knowledge that created the Haber-Bosch process [1] helps to sustain the majority of the world's populous, add another five hundred trillion dollars for that just to start with.

                The creation of the printing press and all written information that allowed it to be built provided dissemination of knowledge beyond the ultra wealthy and is worth a non-finite amount of money.

                LLM's are cool math, but they are less than a rounding error in comparison to even the tiniest sliver of human knowledge and technological output.

                [0] <a href="https:&#x2F;&#x2F;pubmed.ncbi.nlm.nih.gov&#x2F;35143880&#x2F;" rel="nofollow">https:&#x2F;&#x2F;pubmed.ncbi.nlm.nih.gov&#x2F;35143880&#x2F; [1] <a href="https:&#x2F;&#x2F;cen.acs.org&#x2F;food&#x2F;agriculture&#x2F;The-industrialization-Haber-Bosch-process&#x2F;101&#x2F;i26" rel="nofollow">https:&#x2F;&#x2F;cen.acs.org&#x2F;food&#x2F;agriculture&#x2F;The-industrialization-H...

                1. DirkH · · focus · HN ↗
                  Shush, this is Hackernews. We all want to have our egos stoked that our industry is the most important in human history and that tech will transform and save us all. Go away with your historical analysis &#x2F;s
              2. Ar-Curunir · · focus · HN ↗
                What do you think mathematicians were doing for centuries before LLMs?

                And also, humans have been doing a lot more work than just mathematics...

          2. phamilton · · focus · HN ↗
            Simple math:

            A training set of 15 trillion tokens is 10 trillion words.

            A penny a word is cheaper than the cheapest beginner freelance writer.

            That makes a training set of 10 trillion words cost $100B.

            Lots of assumptions there for sure, but we&#x27;re certainly in the ballpark you are describing.

            1. ToValueFunfetti · · focus · HN ↗
              How much did you get paid to write this?
          3. louiskottmann · · focus · HN ↗
            Given they ingested basically the whole internet and then some, you cannot possibly be serious when you mean it&#x27;s worth less than 10 billions.

            The totality of the content on internet is worth several orders of magnitude more.

            1. layer8 · · focus · HN ↗
              The argument wasn’t about how much it’s worth, but about how much it cost to create. These are very different things.
              1. tpm · · focus · HN ↗
                do you also count eg published results of very expensive physics experiments? because once the costs of things like these are taken into account, we are way over 10 billions.
              2. allturtles · · focus · HN ↗
                Why? Are we doing labor theory of value now?
          4. jonhohle · · focus · HN ↗
            If you look at movies alone that would easily surpass 10s of billions. The cost of most books is probably more nebulous, but books, research, and more all have time and money spent to create them. I would guess the corpus of all media from the 20th century on would be minimally in the hundreds of billions of dollars.
            1. layer8 · · focus · HN ↗
              LLMs aren’t trained on movies, though.

              Image&#x2F;video models are, but those weren’t the topic.

          5. Enginerrrd · · focus · HN ↗
            Yes, Easily, and by multiple orders of magnitude.
          6. rsingel · · focus · HN ↗
            I asked Claude to estimate the cumulative salaries of US only journalists over the last hundred years:

            $500B for all kinds including TV and online

            $300B for newsrooms including all staff

            $140B for newsroom reporters only

            So yeah, I think the price of the information ingested is way higher than training costs

          7. Ohentis · · focus · HN ↗
            I suspect so. It is a lot of data. You&#x27;re looking at essentially all publicly available (and some non public) intellectual work.
          8. wonnage · · focus · HN ↗
            this is the sort of brain rot thought that you have in a dorm room the day you are introduced to Econ

            “bro like, what if we could price the sum total of human knowledge? That wouldn’t be that much, right?”

          9. MathiasPius · · focus · HN ↗
            I would argue that producing the complete written corpus on which they at least intend to train (even if some is still out of reach) cost literally everything to produce.

            And the monetary cost doesn&#x27;t even register when weighed against the blood, sweat and tears that went into capturing the authentic experiences of real human beings, whose honest expressions are now at least in some cases getting hoovered up, ingested, and then destroyed for all eternity, for fear that this specific work is the rounding error that might give an equally immoral competitor the edge in the bicycle-riding flamingo race that is currently consuming an absurd amount of the world&#x27;s creativity and attention.

          10. zelphirkalt · · focus · HN ↗
            It probably cost vastly more than training the LLM. You need to consider the time people spent and perhaps weigh it in their hourly wage. Accumulate that across all the training data and it will be a mindboggling amount, compared to training the model.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.