‹ BackHN Continuity

Thread

Can gzip be a language model?

414 points · 165 comments · networked

  1. mg · · focus · HN ↗

        give it a normal text prompt, and it
        continues that prompt by searching
        for the byte sequences that compress
        best.
    
    One moment, how are we supposed to know how well that search was done? There is no way to search a meaningful part of the search space.

    So the result only gives us some lower bound of how well gzip works as a "plausibility tester" of a continuation of a text. The space of possible sequences is many orders of magnitude larger than what was searched. So there might be sequences in there that compress much better.

    The text mentions beamsearch, but I don't see a discussion about how well beamsearch performs in finding the global optima when it comes to gzip compressibility of a text?

    1. StilesCrisis · · focus · HN ↗
      Read to the end: they aren't actually looking for the best-compressing output, because this quickly devolves into aaaaaaaaaaa. They keep a sliding window over a small portion of recent text and use that.

      Basically I think the entire premise falls apart due to that choice--they forced an interesting-looking outcome by adjusting the algorithm until gzip started picking random slabs of letters instead of ever-larger repeating runs.

      1. Ohentis · · focus · HN ↗
        I had actually thought of doing this, but didn't for this exact reason. I knew I would have to fudge things to make it anything interesting.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.