‹ BackHN Continuity

Thread

Can gzip be a language model?

414 points · 165 comments · networked

  1. berkes · · focus · HN ↗
    I've been pondering on something related: can an LLM be a chat?

    Some models are reproducible, in that the same prompt will generate the same output. Say that we could wire up such a model to generate some code.

    In that case, we could create a prompt that generates, say, an entire codebase, or a large piece of text. The prompt (or really, the tokens) would then be the compressed version of the codebase or the text.

    I am not talking about an &quot;AI agent&quot;, but really a model that we call in a reproducible manner. Preferably one call, with one prompt. An agent could just run `git clone` to &quot;decompress&quot; a codebase, which conflates the idea of compression. If that were compression, then the &quot;compressed version of the git kernel&quot; would be a single line of text: `git clone <a href="https:&#x2F;&#x2F;git.kernel.org&#x2F;pub&#x2F;scm&#x2F;linux&#x2F;kernel&#x2F;git&#x2F;torvalds&#x2F;linux.git" rel="nofollow">https:&#x2F;&#x2F;git.kernel.org&#x2F;pub&#x2F;scm&#x2F;linux&#x2F;kernel&#x2F;git&#x2F;torvalds&#x2F;lin...`. I am really talking about having an LLM re-generate text based on a prompt.

    Does that make sense? I can imagine that this is highly impractical and inefficient. But would this count as &quot;compression&quot; at all?

    1. evgpbfhnr · · focus · HN ↗
      You&#x27;re describing <a href="https:&#x2F;&#x2F;bellard.org&#x2F;ts_zip&#x2F;" rel="nofollow">https:&#x2F;&#x2F;bellard.org&#x2F;ts_zip&#x2F; (&quot;Text Compression using Large Language Models&quot;) ?
      1. stackbutterflow · · focus · HN ↗
        Man,did that page load fast. It made me realize how slow the rest the (my) web is.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.