(ass. sports.txt politics.txt and business.txt are text docs pertaining from the sports, politics and business domains, respectively, and have equal size)
The test file belongs to the topic with the smallest size *.gz file.
Witten's group at Waikato uni were perhaps the first to work on this.
Also check out the Hutter prize if you are interested in this.
For anyone wanting an introductory text for information theory & that explores some of these connections & applications, it's worth checking out the late David MacKay's 2003 textbook Information Theory, Inference & Learning Algorithms <a href="https://www.inference.org.uk/itila/" rel="nofollow">https://www.inference.org.uk/itila/
3Blue1Brown has a good video about this: <a href="https://www.youtube.com/watch?v=l6DKRf-fAAM" rel="nofollow">https://www.youtube.com/watch?v=l6DKRf-fAAM
That requires you to decide what a "word" is, which is not trivial (if you think that ignoring punctuation gets you to a clean "letters surrounded by spaces" you will get lots of issues with various Asian languages)
Also some languages have a lot of prefixes and suffixes on their verbs or even nouns, which dilutes your list of 1000 words by just adding the same common words over and over again with different suffixes designating grammatical tense, grammatical gender, etc.
The gzip version sounds more general and more obviously correct
Instead of deciding on words, maybe you can break them up into smaller, subword parts – let's call them "quantums". And these quantums can be the units the quantumizer works on to operate on inputs and outputs. We can then use them to build Expansive Dictionary Models, or EDMs. I suppose we'd need a software library to mak working on this easier, think something speedy, fast, hot, like fire: we can call it PHPFlame...
I'm guessing this is an allusion to reinventing something that already exists, but do you mind explaining what that is to me, since I don't know?
Byte Pair Encoding [1] will be different for different languages. Application of the per-language BPEs to the input text will produce encodings with different lengths.
It naturally takes care of common prefixes and suffixes.
It is easy and fast to apply using radix tree or with finite automata. Even without radix tree, it is possible to have processing speed in the range of hundredths of thousands of bytes per second.
See: Fast Static Symbol Table: <a href="https://github.com/duckdb/duckdb/pull/4366" rel="nofollow">https://github.com/duckdb/duckdb/pull/4366
FSST is based on a fixed size (255 items) dictionary of high frequency variable length strings/substrings (learned from the corpus) encoded as one byte.
Nitpick: Doing it exactly like this is flawed because you let the compressibility of your references taint the result; what you would prefer is the compressed size of testfile given sports.txt/... as a dictionary without accounting for the compressed size of that, no?
You're right, you should subtract off the compressed sizes of the respective reference files before comparing. (This suffices if we assume that later input data does not influence the compression of earlier input data, which is true except for certain unusual conditions like a repeated substring at the end of the reference data that also appears at the beginning of the test data.)
By looking at mutual information from different authors on the same topic vs same author on different topics. As I recall, it convincingly disproved the hypothesis.
You might also want the topic files to be compressed against each other to get a baseline matrix and then multiply any results by the inverse, assuming equal priors on the topics.
jll29 · · focus · HN ↗
The test file belongs to the topic with the smallest size *.gz file.
Witten's group at Waikato uni were perhaps the first to work on this.
Also check out the Hutter prize if you are interested in this.
LPisGood · · focus · HN ↗
Also, I’ve never seen “ass.” Used to shorten “aside” — I typically use N.B. but perhaps only for important ones.
shoo · · focus · HN ↗
arrowsmith · · focus · HN ↗
matzf · · focus · HN ↗
chrisweekly · · focus · HN ↗
stingraycharles · · focus · HN ↗
I seeded gzip compressors’ dictionaries with Wikipedia articles in different languages.
I would then try to use said dictionaries on any random text, and the one that was best able to compress it, was the correct language.
Absolutely totally not the best approach, but very fast and super simple to implement.
actionfromafar · · focus · HN ↗
ape4 · · focus · HN ↗
wongarsu · · focus · HN ↗
Also some languages have a lot of prefixes and suffixes on their verbs or even nouns, which dilutes your list of 1000 words by just adding the same common words over and over again with different suffixes designating grammatical tense, grammatical gender, etc.
The gzip version sounds more general and more obviously correct
basilgohar · · focus · HN ↗
ashkankiani · · focus · HN ↗
BobaFloutist · · focus · HN ↗
dspillett · · focus · HN ↗
giancarlostoro · · focus · HN ↗
kragen · · focus · HN ↗
npb · · focus · HN ↗
thesz · · focus · HN ↗
[1] <a href="https://en.wikipedia.org/wiki/Byte-pair_encoding" rel="nofollow">https://en.wikipedia.org/wiki/Byte-pair_encoding
It naturally takes care of common prefixes and suffixes.
It is easy and fast to apply using radix tree or with finite automata. Even without radix tree, it is possible to have processing speed in the range of hundredths of thousands of bytes per second.
colejohnson66 · · focus · HN ↗
<a href="https://www.youtube.com/watch?v=m8niIHChc1Y" rel="nofollow">https://www.youtube.com/watch?v=m8niIHChc1Y
wodenokoto · · focus · HN ↗
ignoramous · · focus · HN ↗
FSST is based on a fixed size (255 items) dictionary of high frequency variable length strings/substrings (learned from the corpus) encoded as one byte.
woadwarrior01 · · focus · HN ↗
<a href="https://en.wikipedia.org/wiki/Normalized_compression_distance" rel="nofollow">https://en.wikipedia.org/wiki/Normalized_compression_distanc...
m-hodges · · focus · HN ↗
[dead]
myrmidon · · focus · HN ↗
Really interesting approach though.
[deleted] · · focus · HN ↗
[deleted]
akoboldfrying · · focus · HN ↗
Lerc · · focus · HN ↗
ape4 · · focus · HN ↗
chris_va · · focus · HN ↗
<a href="https://en.wikipedia.org/wiki/Baconian_theory_of_Shakespeare_authorship" rel="nofollow">https://en.wikipedia.org/wiki/Baconian_theory_of_Shakespeare...
By looking at mutual information from different authors on the same topic vs same author on different topics. As I recall, it convincingly disproved the hypothesis.
[deleted] · · focus · HN ↗
[deleted]
anthk · · focus · HN ↗
jjtheblunt · · focus · HN ↗
mrtnmcc · · focus · HN ↗