‹ BackHN Continuity

Thread

Can gzip be a language model?

414 points · 165 comments · networked

  1. berkes · · focus · HN ↗
    I've been pondering on something related: can an LLM be a chat?

    Some models are reproducible, in that the same prompt will generate the same output. Say that we could wire up such a model to generate some code.

    In that case, we could create a prompt that generates, say, an entire codebase, or a large piece of text. The prompt (or really, the tokens) would then be the compressed version of the codebase or the text.

    I am not talking about an &quot;AI agent&quot;, but really a model that we call in a reproducible manner. Preferably one call, with one prompt. An agent could just run `git clone` to &quot;decompress&quot; a codebase, which conflates the idea of compression. If that were compression, then the &quot;compressed version of the git kernel&quot; would be a single line of text: `git clone <a href="https:&#x2F;&#x2F;git.kernel.org&#x2F;pub&#x2F;scm&#x2F;linux&#x2F;kernel&#x2F;git&#x2F;torvalds&#x2F;linux.git" rel="nofollow">https:&#x2F;&#x2F;git.kernel.org&#x2F;pub&#x2F;scm&#x2F;linux&#x2F;kernel&#x2F;git&#x2F;torvalds&#x2F;lin...`. I am really talking about having an LLM re-generate text based on a prompt.

    Does that make sense? I can imagine that this is highly impractical and inefficient. But would this count as &quot;compression&quot; at all?

    1. meindnoch · · focus · HN ↗
      &gt;I&#x27;ve been pondering on something related: can an LLM be a chat?

      A chat?

      &gt;I am not talking about an &quot;AI agent&quot;, but really a model that we call in a reproducible manner.

      An LLM is just as deterministic as any other computer program. For identical inputs (which includes the PRNG seed) it produces identical outputs.

      &gt;compressed version of the git kernel

      The git kernel, got it.

      &gt;But would this count as &quot;compression&quot; at all?

      Yes. The decompressor is several tens of gigabytes though.

      1. foldr · · focus · HN ↗
        &gt;An LLM is just as deterministic as any other computer program. For identical inputs (which includes the PRNG seed) it produces identical outputs.

        This is not really true in practice because of multi-threading and out-of-order execution. Mathematically equivalent orderings of operations are not equivalent when dealing with floating point values, so most practical LLM implementations end up being non-deterministic.

        1. MarkusQ · · focus · HN ↗
          &quot;An LLM is just as deterministic as any other computer program&quot; is not refuted by pointing out hardware limitations that would affect any other computer program implemented at similar scale (weather forecasts, or even just computing the average of a large stream of sensor readings).
          1. foldr · · focus · HN ↗
            I&#x27;m not really trying to &#x27;refute&#x27; the original statement. Certainly, an LLM is just doing some calculations that can be done deterministically in principle. However, I think it&#x27;s worth pointing out that there are practical barriers to doing those particular calculations both deterministically and efficiently. People who worry about LLM output not being reproducible aren&#x27;t necessarily misunderstanding what an LLM is doing; they are responding to a real feature of most practical LLM implementations.
            1. eru · · focus · HN ↗
              Agreed.

              If you wanted to and had enough engineering effort to spare, you could run an LLM deterministically at relatively small impacts to performance.

              One approach is to make sure you run things in the same order. Another is to change your operations so that more of them become associative or even commutative.

              See eg the paper &#x27;A Lattice-Based Approach to Deterministic Parallelism&#x27; for some interesting ideas on the latter.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.