‹ BackHN Continuity

Thread

Can gzip be a language model?

414 points · 165 comments · networked

  1. jll29 · · focus · HN ↗
    Yes: you can classify a test file by topic with gzip as follows:

      gzip -9 sports.txt   testfile.txt
    
      gzip -9 politics.txt testfile.txt
    
      gzip -9 business.txt testfile.txt
    
    (ass. sports.txt politics.txt and business.txt are text docs pertaining from the sports, politics and business domains, respectively, and have equal size)

    The test file belongs to the topic with the smallest size *.gz file.

    Witten's group at Waikato uni were perhaps the first to work on this.

    Also check out the Hutter prize if you are interested in this.

    1. myrmidon · · focus · HN ↗
      Nitpick: Doing it exactly like this is flawed because you let the compressibility of your references taint the result; what you would prefer is the compressed size of testfile given sports.txt/... as a dictionary without accounting for the compressed size of that, no?

      Really interesting approach though.

      1. [deleted] · · focus · HN ↗

        [deleted]

      2. akoboldfrying · · focus · HN ↗
        You're right, you should subtract off the compressed sizes of the respective reference files before comparing. (This suffices if we assume that later input data does not influence the compression of earlier input data, which is true except for certain unusual conditions like a repeated substring at the end of the reference data that also appears at the beginning of the test data.)
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.