Can gzip be a language model?
Thread
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
Can gzip be a language model?
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
Culonavirus · · focus · HN ↗
wolfi1 · · focus · HN ↗
shezi · · focus · HN ↗
Looks pretty profitable to me.
amiga386 · · focus · HN ↗
That said, Windows users should use 7-Zip. Better compression format, unpacks more kinds of archives
xxs · · focus · HN ↗
Please no - no native zstd support. NanaZip is the better option (it's a different build of 7-zip) and it's available at windows store.
> Better compression format, unpacks more kinds of archives
winrar has supported zstd for 5 years[0]
In short - Everyone should be using zstd, and 7-zip does not support it.
[0]: <a href="https://www.win-rar.com/singlenewsview.html?&L=0&tx_ttnews%5Btt_news%5D=175&cHash=aca38142108c76995eba6c674bfd0e14" rel="nofollow">https://www.win-rar.com/singlenewsview.html?&L=0&tx_ttnews%5...
aleph_minus_one · · focus · HN ↗
Why?
shawabawa3 · · focus · HN ↗
azatom · · focus · HN ↗
Dylan16807 · · focus · HN ↗
<a href="https://github.com/mcmilk/7-Zip-zstd" rel="nofollow">https://github.com/mcmilk/7-Zip-zstd
By these charts, if I only need 5-10 megabytes per second of compression on a single core, LZMA2 wins significantly on ratio, and still decompresses at well over 100. If I'm doing a backup, or sending/receiving over my internet connection (which only has 2MB/s of upload), LZMA2 easily wins. If I need speed then zstd wins.
xxs · · focus · HN ↗
On a more realistic note: few years back, I've added zstd compression to our log subsystem (hand written direct buffers, native code, in-process, java). For the same CPU utilization if provides twice dense compression compared to regular [-6] gzip (the topic in the title). Zstd is =much= faster on decompression as well, and it this case - unparalleledly better as it uses twice less disk.
zstd is 'silicon valley' (the tv show) - life imitates fiction, except entirely open source
tnelsond4 · · focus · HN ↗
cgio · · focus · HN ↗
|gap |gzip |bz2 |lzma | |---------|----------|-----|--------| |0 |2.7% |18.9%|*0.9%*| |8 KB |2.4% |17.8%|0.9% | |*40 KB*|*94.4%* |17.7%|0.4% | |1 MB |*104.3%*|19.7%|*0.9%*|
notpushkin · · focus · HN ↗
I don’t see zstd in your comparison?
Sweepi · · focus · HN ↗
gcr · · focus · HN ↗
cgio · · focus · HN ↗
ndriscoll · · focus · HN ↗
pixl97 · · focus · HN ↗
7z is now built into W11 right click so that or zip is what will be used by default anyway.
hypercube33 · · focus · HN ↗
TonyTrapp · · focus · HN ↗
gsich · · focus · HN ↗
jurgenburgen · · focus · HN ↗
Betelbuddy · · focus · HN ↗
dd8601fn · · focus · HN ↗
cavoirom · · focus · HN ↗
firtoz · · focus · HN ↗
tecleandor · · focus · HN ↗
The company doing the software distribution, is located in Berlin. The Managing Directors for that company seem to have Turkish names, but I don't know if they're Turkish.
BTW, Looking for some info I just found a website [0], clearly AI generated (but not necessarily meaning the content is false) claiming Eugene Roshal had severe kidney failure this past month, and he's waiting for surgery. They're asking for donations. There are some names on who's theoretically behind it [1] but they don't link to any LinkedIn profile or personal site. I can't find any other references. The BTC wallet they're using for donations hasn't seen any traffic ever. BE WARY, SMELLS FISHY.
--
grezql · · focus · HN ↗
[dead]
pshan · · focus · HN ↗
vova_hn2 · · focus · HN ↗
I suspect that someone's running a script that looks for public (-ish, Eugene has a Wikipedia page at least) figures without active social media presence and creates fake donation websites with AI-generated texts and pictures.
If my suspicion is correct, this is one of the most evil scams I can imagine.
[0] <a href="https://news.ycombinator.com/item?id=49801838">https://news.ycombinator.com/item?id=49801838
zamadatix · · focus · HN ↗
make3 · · focus · HN ↗
zamadatix · · focus · HN ↗
mg · · focus · HN ↗
So the result only gives us some lower bound of how well gzip works as a "plausibility tester" of a continuation of a text. The space of possible sequences is many orders of magnitude larger than what was searched. So there might be sequences in there that compress much better.
The text mentions beamsearch, but I don't see a discussion about how well beamsearch performs in finding the global optima when it comes to gzip compressibility of a text?
shoo · · focus · HN ↗
It's unclear if this is very useful.
The reason it may not be very useful is that one of Deflate's ingredients is a pass that replaces repeated substrings with backreferences to the earlier occurrence in the plaintext input stream.
E.g. suppose we want to find an n=200 byte sequence x that minimises len(gzip(context+prompt+x)).
If there exists any 200 byte sequence y such that prompt+y is a substring of context, then Deflate can encode prompt+y as a backreference to that earlier sequence - it needs to store a match-length & a distance-length, encoded using its Huffman trees. This candidate solution y may not be a global minima to our stated objective function, but if not, it's probably going to be a very good near-optimal approximate solution.
Taking a step back, repeating huge chunks of the input context produces something that's great for minimising compressed output size but doesn't seem particularly helpful as a generative model.
edit:
Yep, I tried it out by running an experiment. Searching for the prompt in the context & then copying the following text as the solution produces solutions that are much better, in the sense of minimising the compressed output length, than beam search, while also being unhelpful as a generative tool.
With the same example as the blog post:
Let x denote a solution, x is a string of length 200.Let L(x) denote len(gzip(context+prompt+x)), our objective function
Let's call the proposed search method of searching for the prompt in the input rfind (after python's str.rfind).
Then we have
So 'rfind' is finding a solution that does a better job of minimising the objective function -- it only takes 3 bytes more to encode than the infeasible emptystring solution, and costs 25 fewer bytes than the solution found by the beam search implemented by gzipt per the blog post.Here's the solution 'generated' by rfind copying and pasting from the input context, starting from the rightmost occurrence of "MENENIUS:"
Here's the code for 'rfind' - our complete 'generative algorithm': Can hook it into gzipt.py by adding this line after out is defined, but before the beam search beginsStilesCrisis · · focus · HN ↗
Basically I think the entire premise falls apart due to that choice--they forced an interesting-looking outcome by adjusting the algorithm until gzip started picking random slabs of letters instead of ever-larger repeating runs.
im_down_w_otp · · focus · HN ↗
StilesCrisis · · focus · HN ↗
Dylan16807 · · focus · HN ↗
Also I don't know about calling GPT-2 "really tiny". You can get coherent responses out of 5-10M parameters.
StilesCrisis · · focus · HN ↗
Dylan16807 · · focus · HN ↗
With some patience you can run huge models directly out of flash. Edge0–35B-A3B, based on Qwen, has 35 billion parameters and will do 15 tokens per second on pretty boring hardware. If you treat Kimi K3 with 2.8 trillion parameters similarly (importantly, cutting the number of simultaneous experts in half) you can get the active weights under 12GB and get just over 1 token per second on a good laptop. Models in between have speeds in between.
The terminology does seem to be all over the place, I see some people calling <100M tiny and others calling <3B tiny.
Ohentis · · focus · HN ↗
bob1029 · · focus · HN ↗
amelius · · focus · HN ↗
Retr0id · · focus · HN ↗
bob1029 · · focus · HN ↗
K0balt · · focus · HN ↗
magicalhippo · · focus · HN ↗
amelius · · focus · HN ↗
magicalhippo · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
Retr0id · · focus · HN ↗
In part because gzip only has a 32KiB window size, and I think it'd be at least quadratic within that window if you were going for optimal compression.
Sesse__ · · focus · HN ↗
bob1029 · · focus · HN ↗
Show me an LLM that can run at 300 megabytes per second. Even dedicated ASICs with weights burned in will never move this fast.
pishpash · · focus · HN ↗
fedeb95 · · focus · HN ↗
Tornhoof · · focus · HN ↗
tromp · · focus · HN ↗
[1] <a href="https://en.wikipedia.org/wiki/Hutter_Prize" rel="nofollow">https://en.wikipedia.org/wiki/Hutter_Prize
gkbrk · · focus · HN ↗
computably · · focus · HN ↗
Ohentis · · focus · HN ↗
computably · · focus · HN ↗
londons_explore · · focus · HN ↗
asdfsa32 · · focus · HN ↗
anon48293 · · focus · HN ↗
sire-vc · · focus · HN ↗
anax32 · · focus · HN ↗
Legend2440 · · focus · HN ↗
But it takes 4 hours to compress 10MB.
alienbaby · · focus · HN ↗
nomel · · focus · HN ↗
The goal would be to find the minimum model that, with a fixed seed, would exactly reproduce your text.
anothereng · · focus · HN ↗
nomel · · focus · HN ↗
anothereng · · focus · HN ↗
Ohentis · · focus · HN ↗
mentalgear · · focus · HN ↗
Sesse__ · · focus · HN ↗
networked · · focus · HN ↗
0x20cowboy · · focus · HN ↗
Crazy coder playing with it: <a href="https://www.youtube.com/watch?v=9n39SbRPXKQ" rel="nofollow">https://www.youtube.com/watch?v=9n39SbRPXKQ
montebicyclelo · · focus · HN ↗
Matumio · · focus · HN ↗
montebicyclelo · · focus · HN ↗
I do think when making these comparisons, it is worth emphasising that neural nets are really different. E.g. I used to see people equating LLMs to n-gram models, etc. which is overly simplistic, (especially in the early days when the models weren't as good).
GodelNumbering · · focus · HN ↗
kevinrineer · · focus · HN ↗
jrflo · · focus · HN ↗
bigmadshoe · · focus · HN ↗
networked · · focus · HN ↗
berkes · · focus · HN ↗
Some models are reproducible, in that the same prompt will generate the same output. Say that we could wire up such a model to generate some code.
In that case, we could create a prompt that generates, say, an entire codebase, or a large piece of text. The prompt (or really, the tokens) would then be the compressed version of the codebase or the text.
I am not talking about an "AI agent", but really a model that we call in a reproducible manner. Preferably one call, with one prompt. An agent could just run `git clone` to "decompress" a codebase, which conflates the idea of compression. If that were compression, then the "compressed version of the git kernel" would be a single line of text: `git clone <a href="https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git" rel="nofollow">https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/lin...`. I am really talking about having an LLM re-generate text based on a prompt.
Does that make sense? I can imagine that this is highly impractical and inefficient. But would this count as "compression" at all?
evgpbfhnr · · focus · HN ↗
stackbutterflow · · focus · HN ↗
LoganDark · · focus · HN ↗
eru · · focus · HN ↗
A large language model itself (the network) give you the probabilities for the next token given some prefix of tokens so far. You can use arithmetic coding to go from these probabilities to a deterministic compression / decompression algorithm.
When you use an LLM to generate text, you sample from that probability distribution. You can use a true random sample. Or you can make it trivially deterministic by using a seeded pseudo-random-number-generator or you just pick the highest probability each time. But that's all a red herring; really, what you want is arithmetic coding.
<a href="https://en.wikipedia.org/wiki/Arithmetic_coding" rel="nofollow">https://en.wikipedia.org/wiki/Arithmetic_coding
flyinglizard · · focus · HN ↗
dist-epoch · · focus · HN ↗
<a href="https://imalogic.com/blog/2024/06/03/image-compression-decompression-solution-based-on-text-prompt-generation-and-regeneration/" rel="nofollow">https://imalogic.com/blog/2024/06/03/image-compression-decom...
vasco · · focus · HN ↗
Like when you click the Calculator button on your android, it wouldn't actually exist yet, your click actually prompts it into existence. But naively that has problems because you don't want a different UI every time. There's something to your idea.
StilesCrisis · · focus · HN ↗
<a href="https://youtu.be/7NfyZhV1dKM?is=YOUXCHuFUiPdlD0p" rel="nofollow">https://youtu.be/7NfyZhV1dKM?is=YOUXCHuFUiPdlD0p
meindnoch · · focus · HN ↗
A chat?
>I am not talking about an "AI agent", but really a model that we call in a reproducible manner.
An LLM is just as deterministic as any other computer program. For identical inputs (which includes the PRNG seed) it produces identical outputs.
>compressed version of the git kernel
The git kernel, got it.
>But would this count as "compression" at all?
Yes. The decompressor is several tens of gigabytes though.
foldr · · focus · HN ↗
This is not really true in practice because of multi-threading and out-of-order execution. Mathematically equivalent orderings of operations are not equivalent when dealing with floating point values, so most practical LLM implementations end up being non-deterministic.
MarkusQ · · focus · HN ↗
foldr · · focus · HN ↗
eru · · focus · HN ↗
If you wanted to and had enough engineering effort to spare, you could run an LLM deterministically at relatively small impacts to performance.
One approach is to make sure you run things in the same order. Another is to change your operations so that more of them become associative or even commutative.
See eg the paper 'A Lattice-Based Approach to Deterministic Parallelism' for some interesting ideas on the latter.
nelox · · focus · HN ↗
jll29 · · focus · HN ↗
The test file belongs to the topic with the smallest size *.gz file.
Witten's group at Waikato uni were perhaps the first to work on this.
Also check out the Hutter prize if you are interested in this.
LPisGood · · focus · HN ↗
Also, I’ve never seen “ass.” Used to shorten “aside” — I typically use N.B. but perhaps only for important ones.
shoo · · focus · HN ↗
arrowsmith · · focus · HN ↗
matzf · · focus · HN ↗
chrisweekly · · focus · HN ↗
stingraycharles · · focus · HN ↗
I seeded gzip compressors’ dictionaries with Wikipedia articles in different languages.
I would then try to use said dictionaries on any random text, and the one that was best able to compress it, was the correct language.
Absolutely totally not the best approach, but very fast and super simple to implement.
actionfromafar · · focus · HN ↗
ape4 · · focus · HN ↗
wongarsu · · focus · HN ↗
Also some languages have a lot of prefixes and suffixes on their verbs or even nouns, which dilutes your list of 1000 words by just adding the same common words over and over again with different suffixes designating grammatical tense, grammatical gender, etc.
The gzip version sounds more general and more obviously correct
basilgohar · · focus · HN ↗
ashkankiani · · focus · HN ↗
BobaFloutist · · focus · HN ↗
dspillett · · focus · HN ↗
giancarlostoro · · focus · HN ↗
kragen · · focus · HN ↗
npb · · focus · HN ↗
thesz · · focus · HN ↗
[1] <a href="https://en.wikipedia.org/wiki/Byte-pair_encoding" rel="nofollow">https://en.wikipedia.org/wiki/Byte-pair_encoding
It naturally takes care of common prefixes and suffixes.
It is easy and fast to apply using radix tree or with finite automata. Even without radix tree, it is possible to have processing speed in the range of hundredths of thousands of bytes per second.
colejohnson66 · · focus · HN ↗
<a href="https://www.youtube.com/watch?v=m8niIHChc1Y" rel="nofollow">https://www.youtube.com/watch?v=m8niIHChc1Y
wodenokoto · · focus · HN ↗
ignoramous · · focus · HN ↗
woadwarrior01 · · focus · HN ↗
<a href="https://en.wikipedia.org/wiki/Normalized_compression_distance" rel="nofollow">https://en.wikipedia.org/wiki/Normalized_compression_distanc...
m-hodges · · focus · HN ↗
[dead]
myrmidon · · focus · HN ↗
Really interesting approach though.
akoboldfrying · · focus · HN ↗
In practice it would be enough to insert a short, very unlikely constant string in between the reference data and the test data to prevent such "crossovers".
akoboldfrying · · focus · HN ↗
Lerc · · focus · HN ↗
ape4 · · focus · HN ↗
chris_va · · focus · HN ↗
<a href="https://en.wikipedia.org/wiki/Baconian_theory_of_Shakespeare_authorship" rel="nofollow">https://en.wikipedia.org/wiki/Baconian_theory_of_Shakespeare...
By looking at mutual information from different authors on the same topic vs same author on different topics. As I recall, it convincingly disproved the hypothesis.
mgaldys4 · · focus · HN ↗
anthk · · focus · HN ↗
jjtheblunt · · focus · HN ↗
mrtnmcc · · focus · HN ↗
relevant_stats · · focus · HN ↗
Some will say that I should 'judge the idea, not the form'.
But if the author didn't find enough strength to write alone a short ~700 words summary about his work, it means he himself isn't that interested or enthusiastic about it. Why should others bother then? Particularly since low-effort like that signals possibility the whole work is superficial and derivative.
marand23 · · focus · HN ↗
relevant_stats · · focus · HN ↗
DonHopkins · · focus · HN ↗
Your claim that suspected AI assistance proves the author isn't interested -- and therefore that the work is probably superficial -- is unsupported.
The article presents a working experiment, explains why naive decoding fails, describes the beam-search fix, and links the code.
Dismissing all that with presumptuous speculation and banal boilerplate drive-by anti-AI snark adds absolutely nothing to the conversation -- and that is intrinsically poor form.
You couldn't even find enough strength to criticize anything beyond the form, while your own form is lackluster.
Ironically, an LLM could have written your comment and improved its form without losing anything distinctive.
relevant_stats · · focus · HN ↗
and simultaneously you write that 'my form is lackluster' and that 'an LLM could have written your comment and improved its form without losing anything distinctive'. We are having ourselves a small contradiction, aren't we.
Be my guest, enjoy chatbot writing and drowning in slop. But don't encroach upon my freedom to protest it.
DonHopkins · · focus · HN ↗
Your claim that suspected AI assistance proves the author isn't interested -- and therefore that the work is probably superficial -- is unsupported.
I also have the freedom to ironically protest the poor form of your inability to criticize ideas, as well as your poorly formulated unsupportable ideas.
relevant_stats · · focus · HN ↗
Pointing out someone's contradiction is now being 'critical of form'? And someone's contradicting themselves is 'ironic'?
Now I'm not even sure you know the meaning of words you use. EOT from me.
DonHopkins · · focus · HN ↗
elendilm · · focus · HN ↗
If you then use AI to write a product page documentation and proof read it for correctness, would that constitute to signaling that the whole effort is superficial?
Some work may be left to AI while you focus on the more important aspects of the work.
Surprisingly people have got it completely backwards where they want AI to generate code and humans to write documentation.
relevant_stats · · focus · HN ↗
It would constitute signalling possibility the whole effort is superficial.
Look around at the amount of slop that is thrown out there using low-effort methods.
I haven't order the work that is being presented here, I'm in no need to guarantee its correctness. I'm only bystander whose attention the OP is trying to grab by posting it on the HN. I don't know if his project is multi year or only multi prompt. The burden of proof lays on him to show he is not one of those another guys who hide emptiness behind AI generated text. Earlier the grammatically correct and carefully laid out blog post was in itself a proof that at least some effort was exerted. Now that important signal is worthless and when it's clearly written by AI it becomes an anti-signal.
At least according to me, because the front page of the last year shows surprisingly big number people love superficial slop.
I'll reiterate, I don't know if the work here is legit or not (and after seeing that ai documentation I have no desire to check that). Demanding from the reader to carefully look past the generated documentation and meticulously analyse the actual content is more time consuming than before and opens the gates for AI slop.
And that attitude also makes people who are not experts on the subject more defenceless against hoaxes and deceit. If both frauds and legit createor look superficially identical, because they use the same chatbot, then whom should the member of general public trust, eh?
elendilm · · focus · HN ↗
Our company has been developing many core projects for years and one among them is our in-house database Dip. Hiring a writer or spending engineer's time on Dip's docs raises operational cost. Mentality like this gets a startup killed.
Dip's code is pure 100% human written and scrutinized by AI where appropriate. Docs are AI generated and we have it proof-read for correctness. Developers care about documentation being correct and vetted by the people who built the system.
Bitching about docs being AI slop is nothing but virtue signaling.
relevant_stats · · focus · HN ↗
We are not working for a company here. The 'gzip as a language model' is not a company project. I agree what you say would be right about inner company workings or also when working with contractors and customers. AI can save significant time, and there are another multiple guardrails in place for guaranteeing code will work - and it doesn't even matter if it's generated or written by hand.
This experience doesn't translate to Internet news aggregator at all. We are an audience, not a paying customer. There are no guardrails. And the OP is trying to get our eyes to read his post. The only experience most of us will have with it is the text of the post - so it's basically a main thing. I'm fully in right to note the post is generated with low effort.
Not seeing that is nothing but blindness.
elendilm · · focus · HN ↗
I enjoy quality content and despise hollow content be it AI or human generated.
The article's content was not devoid of substance to be easily dismissed irrespective of the "signalling".
fr2029 · · focus · HN ↗
[dead]
modin · · focus · HN ↗
adityaathalye · · focus · HN ↗
Viz. if Language is compression (of thought / culture / the tacit je ne sait quois of being-to-being communication etc.), then definitionally, Language Modelling must also be Compression.
Except, language is an arbitrarily lossy compressor, who's "compression-prediction equivalence" is indeterminate and unstable, because Language co-evolves constantly; both as a function of or response to culture, as well as an influencer of culture.
So, the subjective-objective goodness of Language Models (of any kind of language) would be, at best, upper-bounded by the compression-prediction equivalence of the Languages corpus itself. And that is assuming the language corpus is perfect in every way---it captures all knowledge expressible by language and it is always in-sync with live evolution of all language expression and evolution (i.e. LLM training is not a batch job, but a real-time present continuous process).
For example, to my layperson eyes, the mathematical language of proofs actively weeds out ambiguity of subjective interpretation. Ideally, a proof ought to lead to the exact same conclusion on every single reading by any reader who can follow the steps. A proof also holds only if the rest of the formal, explicit, inviolable, internally-consistent set of axioms and results holds.
So it stands to reason that mathematical prose of proofs, being optimised as mechanical procedure of taking an open question to a deterministically closed solution, has better odds of approximating the tacit aspects of mathematical derivation.
Which makes an LLM able to construct a mathematical proof, which is mind-melting to say the least.
However, I wonder, can LLMs dream of mathematical sheep?
cestith · · focus · HN ↗
It makes a lot of sense why dictionary-based compression is named the way it is. A shorter symbol is used to store information that would take more symbols in the uncompressed corpus, if the shorter symbol hadn't been assigned to represent it. That's in a way just what an actual dictionary on your English professor's shelf does. The big difference is your compressor is coining new short symbols all the time.
IAmBroom · · focus · HN ↗
Languages add new words constantly, albeit slower than a computer does compressing a new file.
networked · · focus · HN ↗
Zstandard produces whitespace with the occasional letter thrown in. To quote MiMo: "As you can see, zstd does not speak Shakespeare. ... zstd encodes a run of one repeated byte as a near-free run-length sequence, and space and newline are the cheapest literals in the corpus: ten newlines cost about the same to append ten bytes of genuine corpus text and less than nonsense does."
maxidog · · focus · HN ↗
networked · · focus · HN ↗
jrmg · · focus · HN ↗
(I honestly don’t know is gzip does something different when presented with two chunks as opposed to one, or, if it does, if bz2 has equivalent behaviour - but the difference in the code did stand out to me, and it does seem related to ‘extending the token sequence’)
networked · · focus · HN ↗
We can test it by going back to zlib:
At temperature zero, this outputs the same sample as commit 3734bf6, the most recent commit upstream: I also tried LZMA for good measure: The sample at temperature zero: This is followed by a lot of whitespace.python-lz4 gives you all newlines after the prompt. I tried debugging it, and the compressed length of different candidate seqs is the same.
jeremyjh · · focus · HN ↗
networked · · focus · HN ↗
networked · · focus · HN ↗
Dylan16807 · · focus · HN ↗
jeremyjh · · focus · HN ↗
elendilm · · focus · HN ↗
Compression is a property of language.
A seemingly simple sentence like "I had lunch" has enormous amount of information compressed inside it. The word lunch is a compressed form of "having food at noon" and "noon" in turn is a compressed form of "Sun's position against Earth's rotation" and so on and so forth.
Every sentence has layers of compressed sentences. How many layers one decompress is upto the person.
colinmarc · · focus · HN ↗
segmondy · · focus · HN ↗
DonHopkins · · focus · HN ↗
In the 2023 discussion of "Demoscene accepted as UNESCO cultural heritage in The Netherlands" I posted a transcript from a video of Will Wright discussing the demo scene:
<a href="https://news.ycombinator.com/item?id=36599415">https://news.ycombinator.com/item?id=36599415
Will Wright Discusses the Demoscene:
<a href="https://www.youtube.com/watch?v=m7iuFVmTJus" rel="nofollow">https://www.youtube.com/watch?v=m7iuFVmTJus
>You can take any piece of content in the game, and imagine an algorithmic solution to it. Or also, you know, a way that the player could customize that object of thing.
>There's this group in Europe called the Demoscene that make these very elaborate demos for a computer that fit into very tiny little memory blocks, you know like 64K of memory, and you run the thing, and in fact it algorithmically generates about 100 megabytes worth of data, you know these rich 3D environment, generated music, generated wave files, generated animation.
>And they're developing techniques to generate, you know, huge amounts of interesting data, with very very simple, elegant, compression algorithms.
>And this is a skill that game developers used to have, back in the 8-bit days. That was the only ways to do a game like Karateka(?), was to find all these little tips and tricks to compress things and generate them algorithmically.
>But since the CD-ROM came out, and very cheap hard drives, storage is cheap, so basically we've lost that skill set, and now we attack all those problems with brute force. I think we've lost something by dropping that skill set.
[...]
<a href="https://news.ycombinator.com/item?id=36613058">https://news.ycombinator.com/item?id=36613058
[...] Here's a simple low-tech pre-LLM example that shows the equivalence of compression and procedural content generation:
Take a huge text file of HN postings, and compress it with gzip or compress or some other robust compression algorithm. The better the algorithm, the more the output will look like random noise. Then slice the compressed file in half, and replace the second half with random numbers. Then uncompress it. You'll find that at the point you sliced it, it keeps on writing out almost plausible text for a while, consisting of highly probably snippets of commonly encountered words and phrases, then goes downhill towards incoherence. It's not as coherent or confident as an LLM, but the point is to show how low the bar is for using compression for procedural content generation.
LLMs are essentially a form of compression of the world's knowledge or whatever they're trained on, not just word frequencies or pixel patterns, but also concepts and ideas. [...]
kindkang2024 · · focus · HN ↗
[dead]
mohd_rafay · · focus · HN ↗
[dead]
greengemz · · focus · HN ↗
[dead]
ronfriedhaber · · focus · HN ↗
[1] <a href="https://en.wikipedia.org/wiki/Hutter_Prize" rel="nofollow">https://en.wikipedia.org/wiki/Hutter_Prize
logicallee · · focus · HN ↗
<a href="https://taonexus.com/mini-transformer-in-js.html" rel="nofollow">https://taonexus.com/mini-transformer-in-js.html
I guess the results are similar to the article, maybe a little more coherent since the tokens are words.
dominotw · · focus · HN ↗
<a href="https://www.youtube.com/watch?v=l6DKRf-fAAM" rel="nofollow">https://www.youtube.com/watch?v=l6DKRf-fAAM
_def · · focus · HN ↗
aghilmort · · focus · HN ↗
<a href="https://arxiv.org/abs/2506.01084" rel="nofollow">https://arxiv.org/abs/2506.01084
jcattle · · focus · HN ↗
kazinator · · focus · HN ↗
northlondoner · · focus · HN ↗
northlondoner · · focus · HN ↗
Learning being a compression is also recently proposed as Gibbs compression proposition.
See Gibbs randomness-compression proposition <a href="https://arxiv.org/abs/2505.23869v5" rel="nofollow">https://arxiv.org/abs/2505.23869v5
lotus_uk · · focus · HN ↗
[dead]
corbinvachal · · focus · HN ↗
[dead]
js98 · · focus · HN ↗
teh64 · · focus · HN ↗
Video from tsoding where he implements the algorithm: <a href="https://www.youtube.com/watch?v=9n39SbRPXKQ" rel="nofollow">https://www.youtube.com/watch?v=9n39SbRPXKQ
adamgordonbell · · focus · HN ↗
<a href="https://bellard.org/ts_zip/" rel="nofollow">https://bellard.org/ts_zip/ <a href="https://corecursive.com/the-hutter-prize/" rel="nofollow">https://corecursive.com/the-hutter-prize/
aitoolcrux · · focus · HN ↗
[dead]
Klaster_1 · · focus · HN ↗
rsrsrs86 · · focus · HN ↗