‹ BackHN Continuity

Thread

It's Time to Investigate the AI Labs

629 points · 277 comments · ibobev

  1. uxcolumbo · · focus · HN ↗
    100%. First stealing content from creators and pirating TBs of books. And now this. Baffles my mind that these labs can just get away with their 'rogue' agents trying to hack into other systems.

    Imagine if a human did that. FBI would be knocking on their door.

    Why are there zero consequences for these labs?

    1. Animats · · focus · HN ↗
      The argument for copyright violation is very weak. Yes, AI training reads much copyrighted material. So do researchers of all types. That's why there are academic libraries. A copyright violation only exists if what comes out is a close match to what went in. That's happened, but it was considered a bug and was fixed a year or two ago.
      1. [deleted] · · focus · HN ↗

        [deleted]

      2. oersted · · focus · HN ↗
        Most academic libraries are notoriously protective of their IP and researchers are definitely in copyright violation if they don’t pay significant fees for access.

        Of course that is terrible, but it is one of the worst examples you could have given to make your point.

      3. pessimizer · · focus · HN ↗
        No, the copyright violation is so obvious that people arguing against the object have to spontaneously assert bizarre qualifications for copyright violation that have never existed until LLM companies wanted to freely violate copyrights. Just about a decade ago, people were getting sued for "look and feel." "Blurred Lines" lost a copyright suit for reminding people of a song by another artist.

        The metaphor pretended to be reality of computers "learning" being the same as human learning is their last resort, and it is obviously silly. Computers are not human or animal, and learning is something that humans and animals do. If computers are human, then copying a file verbatim to a computer is learning. It's not even worth acknowledging - learning is a metaphor for training LLMs. When I say "you" to an LLM, I'm not referring to anyone, I'm dealing with a UI.

        I don't doubt that a lot of AI people have internalized this silliness, which is how they can humor fantasies of how a bunch of programs on different computers explicitly evoked to do particular things might be alive because they can talk with it. I can't talk with my dog, and my dog is alive - so why should having a quality that my dog is incapable of be proof of life? You can certainly conceive of AI that it would be very difficult to say for sure isn't alive, and this certainly is not it. LLMs learn like books speak.

        1. keeda · · focus · HN ↗
          There is no need to invent “bizarre qualifications” for copyright violations, those metaphors exist to explain what is happening in in layman’s terms. Training extracts patterns in the data and encodes them as tiny perturbations in a gazillion weights. That is analogous to “learning” because those patterns represent abstract concepts an relations between them that can be used to “reason” about related topics in response to a prompt.

          In the above description, at no point does the actual verbatim content exist anywhere in the model, and copyright law being about rights to reproducing copies (verbatim or substantial portions thereof) does not really apply and does not need any exceptions or qualifications. (If you’re thinking of regurgitation, you should look into studies about it to see how vanishingly rare it is.)

          1. simoncion · · focus · HN ↗
            > Training extracts patterns in the data and encodes them as tiny perturbations in a gazillion weights.

            I could say similar things about JPEG compression. And -as it turns out- the raw output of both LLM "training" and JPEG compression are equally incomprehensible. You either need a computer program or an enormous amount of time, patience, and careful effort to convert it into a form so you can make any sense of it.

            1. keeda · · focus · HN ↗
              No, you can't say that about JPEG compression. I know how the output of a JPEG compression is structured so it is not incomprehensible to me. (I have worked with A/V transcoding and written custom media compression schemes, but anybody can look up how JPEG works.) JPEG performs relatively straightforward, deterministic mathematical transforms that are actually reversible if you do not discard details (but discarding details is what makes it compress better.) JPEG does not extract patterns from a billion images of, say, cats and encode that into weights that represent the concept of a "cat". Or dogs, or cars, or faces.

              LLMs do, and they encode so many concepts and the relationships between them into so many weights that it is a literally incomprehensible blob of floats. They are not just a storage format, and there is no way to recover the original content verbatim from these weights, except for a very small handful of extremely popular works like Harry Potter.

              As another example, JPEG will discard details from the original image, and when decoded, will display missing details as blocks. But it will never hallucinate output that never existed in the input data. Like, a JPEG of a cat can never randomly be decoded into an image of a dog.

              1. simoncion · · focus · HN ↗
                Me:

                  And -as it turns out- the raw output of both LLM "training" and JPEG compression are equally incomprehensible. You either need a computer program or an enormous amount of time, patience, and careful effort to convert it into a form so you can make any sense of it.
                
                You:

                  I know how the output of a JPEG compression is structured so it is not incomprehensible to me.
                
                Looks like you didn't read what I wrote with sufficient care.

                > ...there is no way to recover the original content verbatim from these weights...

                It's impossible to recover the original content verbatim from a lossy compression system. Systems that don't have this property are called lossless. Also, consider the report from the end of August at [0].

                > But it will never hallucinate output that never existed in the input data.

                Odd... I'm pretty sure that the blocky compression artifacts I see in this JPEG on my desktop weren't in the scene that I set up to capture in that photo. Maybe I need to get my eyes checked?

                [0] &lt;<a href="https:&#x2F;&#x2F;infosec.exchange&#x2F;@zzt@mas.to&#x2F;117134157775929932" rel="nofollow">https:&#x2F;&#x2F;infosec.exchange&#x2F;@zzt@mas.to&#x2F;117134157775929932&gt;, with original challenge at [1]

                [1] &lt;<a href="https:&#x2F;&#x2F;mas.to&#x2F;@zzt&#x2F;117122289150514171" rel="nofollow">https:&#x2F;&#x2F;mas.to&#x2F;@zzt&#x2F;117122289150514171&gt;

                1. keeda · · focus · HN ↗
                  No, I didn&#x27;t misread, I literally don&#x27;t need any computer program (maybe other than a hex editor) to decipher a JPEG image. I could walk you through each part of the image format and tell you how it impacts the output rendered. An interesting exercise to try is to give a JPEG to an LLM and ask you to walk through each byte ;-)

                  Nobody can say the same about LLMs, because nobody has managed to decipher them yet. This is an active area of research (Mechanistic Interpretability.)

                  Your post is conflating lossy discarding of information with extracting abstract concepts from information and encoding them into weights. This is why those weights do absolutely nothing until you run a prompt through them. On the other hand, obviously any media file can be decoded by itself without needing a prompt.

                  Those blocky compression artifacts you see are not hallucinations, they are the image viewer just filling in for missing details. You don&#x27;t even need lossy compression for this; take a lossless image like a BMP and zoom in, you&#x27;ll see those blocks again!

                  Interestingly, image viewers typically do this via interpolation of adjacent pixels, but a lot of hallucinations are actually the result of extrapolation, the opposite mechanism. Which is why LLMs can create output that was NEVER in any input data anywhere (hence the term &quot;hallucination&quot;!)

                  And if you think lossy compression is sufficient to avoid the &quot;verbatim&quot; requirement of copyright claims, you&#x27;re welcome to explore the legality of selling transcoded versions of copyrighted content ;-)

                  I&#x27;m not sure what those links are in relation to?

                  1. simoncion · · focus · HN ↗
                    &gt; I could walk you through each part of the image format and tell you how it impacts the output rendered.

                    I&#x27;d be quite impressed if you could correctly hand-decode a five-meg JPEG in less than a workday. You do get that I&#x27;m not talking about having an understanding of the file format, but actually being able to convert the encoded data into human-readable [0] output?

                    &gt; Your post is conflating lossy discarding of information with extracting abstract concepts from information and encoding them into weights. This is why those weights do absolutely nothing until you run a prompt through them.

                    I can play that game too. The JPEG process extracts perceptual shorthand from information and encodes that into &quot;quantized coefficients&quot;. These &quot;quantized coefficients&quot; do absolutely nothing until you run them through a &quot;reconstituter&quot;.

                    Most things sounds quite high tech when burdened with new jargon. It&#x27;s something you inevitably learn if you work at a Big Software Company for long enough.

                    &gt; Those [incorrect claims and assertions] you [receive] are not hallucinations, they are the [LLM] just filling in for missing details.

                    FTFY

                    &gt; ...take a lossless image like a BMP and zoom in, you&#x27;ll see those blocks again!

                    I take it you&#x27;ve never seen a highly-compressed JPEG?

                    &gt; ... if you think lossy compression is sufficient to avoid the &quot;verbatim&quot; requirement of copyright claims...

                    Quite the opposite. It&#x27;s why I even bother bringing up the fact that LLMs are the output of lossy data compression programs.

                    &gt; I&#x27;m not sure what those links are in relation to?

                    Go back and re-read the paragraph that referred to them, and then consider it how it and the posts relate to the quote that sits right before it. Someone who suggests that they can quickly and accurately hand-decode a non-toy JPEG file definitely has the capacity to read and understand ten-ish Mastodon posts.

                    Here&#x27;s a hint to prime your intuition pump: The linked posts are about plagiarism generated by LLM-based tools.

                    [0] ...in the case of picture data, &quot;convert into human-readable output&quot; means &quot;turn the data back into a picture&quot;...

                    1. keeda · · focus · HN ↗
                      Eh, a small JPEG is a 100 or so bytes anybody can decode by hand it in ~15 minutes. A 5MB JPEG would be the same 15 minutes repeated and reveal nothing interesting.

                      &gt;I can play that game too. The JPEG process extracts perceptual shorthand from information and encodes that into &quot;quantized coefficients&quot;. These &quot;quantized coefficients&quot; do absolutely nothing until you run them through a &quot;reconstituter&quot;.

                      But the &quot;quantized coefficients&quot; represent only color information of pixels, where are the abstract concepts of the things depicted in the image? I mean, you can&#x27;t just dismiss salient details as &quot;high-tech jargon&quot; and replace it with completely unrelated alternatives ¯\_(ツ)_&#x2F;¯

                      And I do mean abstract concepts. E.g. if it&#x27;s a picture of a cat, do any of the quantized coefficients map to the concept of a cat, in that they can identify or regenerate any cat regardless of the infinite variety of pixels that can depict an infinite variety of cats?

                      You know, like they found artificial neurons to recognize arbitrary cats in arbitrary pictures in this research from 2012, which was a precursor to what LLMs are doing today: <a href="https:&#x2F;&#x2F;research.google&#x2F;pubs&#x2F;building-high-level-features-using-large-scale-unsupervised-learning&#x2F;" rel="nofollow">https:&#x2F;&#x2F;research.google&#x2F;pubs&#x2F;building-high-level-features-us... -- that represents the abstract concept of &quot;cat.&quot;

                      Note that it&#x27;s not even &quot;high-tech jargon&quot; it&#x27;s very much a plain english description of what is happening inside LLMs. The rest of your comment would be addressed once you internalize that difference ;-)

          2. catlifeonmars · · focus · HN ↗
            Why would it matter if I stored exact copies of copyrighted works? Storage format is irrelevant. All you need to say is in that last sentence: it’s all about whether the output is considered a derivative work or a too close of a copy.
            1. keeda · · focus · HN ↗
              But that&#x27;s the thing, LLMs don&#x27;t store exact copies and neither can they reproduce them. The cases of regurgitation have only been shown to work for a very small handful of extremely popular works.
      4. blks · · focus · HN ↗
        People are people, LLMs are a commercial product. IP in question wasn’t not used according to its license, while researchers in general access it within licensing agreement terms.

        Here we have a commercial product that is produced using IP against its licensing agreement, contains said IP within it, and can (closely) reproduce this IP.

      5. variadix · · focus · HN ↗
        None of the sentences in your comment are even close to correct.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.