‹ BackHN Continuity

Thread

Can gzip be a language model?

414 points · 165 comments · networked

  1. jll29 · · focus · HN ↗
    Yes: you can classify a test file by topic with gzip as follows:

      gzip -9 sports.txt   testfile.txt
    
      gzip -9 politics.txt testfile.txt
    
      gzip -9 business.txt testfile.txt
    
    (ass. sports.txt politics.txt and business.txt are text docs pertaining from the sports, politics and business domains, respectively, and have equal size)

    The test file belongs to the topic with the smallest size *.gz file.

    Witten's group at Waikato uni were perhaps the first to work on this.

    Also check out the Hutter prize if you are interested in this.

    1. stingraycharles · · focus · HN ↗
      Back in the day - maybe two decades ago - I implemented language detection like this.

      I seeded gzip compressors’ dictionaries with Wikipedia articles in different languages.

      I would then try to use said dictionaries on any random text, and the one that was best able to compress it, was the correct language.

      Absolutely totally not the best approach, but very fast and super simple to implement.

      1. ape4 · · focus · HN ↗
        Or maybe make a list of the most used 1000 words in each language. And see which list has the most occurrences.
        1. wongarsu · · focus · HN ↗
          That requires you to decide what a "word" is, which is not trivial (if you think that ignoring punctuation gets you to a clean "letters surrounded by spaces" you will get lots of issues with various Asian languages)

          Also some languages have a lot of prefixes and suffixes on their verbs or even nouns, which dilutes your list of 1000 words by just adding the same common words over and over again with different suffixes designating grammatical tense, grammatical gender, etc.

          The gzip version sounds more general and more obviously correct

          1. basilgohar · · focus · HN ↗
            Instead of deciding on words, maybe you can break them up into smaller, subword parts – let's call them "quantums". And these quantums can be the units the quantumizer works on to operate on inputs and outputs. We can then use them to build Expansive Dictionary Models, or EDMs. I suppose we'd need a software library to mak working on this easier, think something speedy, fast, hot, like fire: we can call it PHPFlame...
            1. ashkankiani · · focus · HN ↗
              I'm guessing this is an allusion to reinventing something that already exists, but do you mind explaining what that is to me, since I don't know?
              1. BobaFloutist · · focus · HN ↗
                I suspect they're talking about LLM tokens.
              2. dspillett · · focus · HN ↗
                It is describing essentially how tokenisation is done for LLMs (Large Language Model → Expansive Dictionary Model).
                1. giancarlostoro · · focus · HN ↗
                  You mean Quantumnisation ;)
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.