‹ BackHN Continuity

Thread

Can gzip be a language model?

414 points · 165 comments · networked

  1. mg · · focus · HN ↗

        give it a normal text prompt, and it
        continues that prompt by searching
        for the byte sequences that compress
        best.
    
    One moment, how are we supposed to know how well that search was done? There is no way to search a meaningful part of the search space.

    So the result only gives us some lower bound of how well gzip works as a "plausibility tester" of a continuation of a text. The space of possible sequences is many orders of magnitude larger than what was searched. So there might be sequences in there that compress much better.

    The text mentions beamsearch, but I don't see a discussion about how well beamsearch performs in finding the global optima when it comes to gzip compressibility of a text?

    1. StilesCrisis · · focus · HN ↗
      Read to the end: they aren't actually looking for the best-compressing output, because this quickly devolves into aaaaaaaaaaa. They keep a sliding window over a small portion of recent text and use that.

      Basically I think the entire premise falls apart due to that choice--they forced an interesting-looking outcome by adjusting the algorithm until gzip started picking random slabs of letters instead of ever-larger repeating runs.

      1. im_down_w_otp · · focus · HN ↗
        Meh. That’s nothing compared to the amount of curation and tuning the LLMs are coerced with.
        1. StilesCrisis · · focus · HN ↗
          No--even a really tiny, underpowered model like GPT-2 with no system prompt produces coherent (though not necessarily desirable or correct) responses.
          1. Dylan16807 · · focus · HN ↗
            That's still an algorithm that underwent immense amounts of tuning. The tuning here on gzip is very simple, very few parameters, and generic. There is no reason to reject it.

            Also I don't know about calling GPT-2 "really tiny". You can get coherent responses out of 5-10M parameters.

            1. StilesCrisis · · focus · HN ↗
              Heck, a few weeks ago someone posted a 25K-ish generator running on an 8-bit machine (ZX Spectrum? BBC Micro?). It emitted semi-plausible random English phrases. So tiny is subjective but I still think GPT-2 qualifies as a tiny model by any modern standard. I could run it on a laptop.
              1. Dylan16807 · · focus · HN ↗
                Modern laptops and desktops are so powerful though. That's not a great way to measure tiny.

                With some patience you can run huge models directly out of flash. Edge0–35B-A3B, based on Qwen, has 35 billion parameters and will do 15 tokens per second on pretty boring hardware. If you treat Kimi K3 with 2.8 trillion parameters similarly (importantly, cutting the number of simultaneous experts in half) you can get the active weights under 12GB and get just over 1 token per second on a good laptop. Models in between have speeds in between.

                The terminology does seem to be all over the place. GPT-2 is definitely a "small" LLM at best, but definitions of "tiny" might cap out at 100M or at 3B...

                I think if I wanted to involve computer capability, I would say tiny is what you can train from scratch on a laptop. Not just run.

      2. Ohentis · · focus · HN ↗
        I had actually thought of doing this, but didn't for this exact reason. I knew I would have to fudge things to make it anything interesting.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.