‹ BackHN Continuity

Thread

I don't want to read what you didn't write

1070 points · 460 comments · mooreds

  1. hatthew · · focus · HN ↗
    As I have been saying for years:

    Writing is fundamentally the transfer of information from your brain to my brain. If you have 1000 bits of semantic information you want to transfer, you can't give 300 bits of semantic information to an LLM and have it fill in the remaining 700, because it doesn't know what those 700 bits are. If it's able to guess those 700 bits correctly, then they aren't true semantic information, and you really only have 300 bits you want to transfer. You might as well transfer those bits to me directly, rather than having the LLM add on an extra superfluous 700 bits that I then have to filter out.

    1. zefalt · · focus · HN ↗
      I’m not sure the 300-bit → 1,000-bit framing applies in all instances. The 300 bits may be a compressed cue to a much fuller idea. The AI can combine that cue with its prior knowledge to help reconstruct what the prompter was trying to express, with the prompter then verifying whether it’s right. Without the relevant prior knowledge for reconstruction, or the prompter for verification, it becomes much harder to know whether you’ve reconstructed the intended idea.
      1. shiandow · · focus · HN ↗
        Unless that 700 bit was transferred on a separate occasion the inferred 700 bits is not true information, anyone could have reconstructed it from the 300 bits.
        1. TeMPOraL · · focus · HN ↗
          No, the 700 bits come from the sender verifying and vouching for the information before sending.
          1. shiandow · · focus · HN ↗
            Unless that takes 700 attempts on average I don't think that actually works.
            1. ben_w · · focus · HN ↗
              That assumes each bit is a coinflip, doesn't it?

              Even Markov chain autocorrect tools do better than 50% odds*, and even GPT-2 was significantly better than that kind of autocorrect.

              * at the word level; IDK how redundant/efficient language is when it comes to bits-worth-of-fact-claims-per-word. But "your cat is sitting on my" -> [mat, laundry, roof, head, belly, laptop, microwave, …] clearly has many bits of information, and a Markov chain will encode the most likely next word even if the user doesn't know what the most likely next word is. Verifying where the cat is sitting is also very easy, as is correction.

              1. hatthew · · focus · HN ↗
                In information theory, each bit is a coin flip by definition
                1. ben_w · · focus · HN ↗
                  I realise I phrased this poorly.

                  I will try harder. Consider entropy.

                  The first sentence in this comment contains 32 characters; from the point of view of a naïve channel with no compression, that's 256 bits (given none require breaking out of the first bytes of UTF-8).

                  It did not take 2^256 attempts to construct the first sentence in this comment, because the generation process was not flipping coins per bit.

                  LLMs also do not emit bits chosen with a [0: 0.5, 1: 0.5] probability distribution.

                  From a compression point of view, the bits-transmitted-per-bits-in-message ratio can be reduced such that more likely messages use fewer bits than less likely messages. However, this requires the receiver to agree with the sender what the probability distribution over tokens is.

                  Intelligence is, amongst other things, a compression algorithm. If I can predict your next token, and we both know this, we can agree in advance that you don't need to actually send it.

                  No single human brain is able to predict the output of an LLM anything like well enough to do that.

                  In entropy terms: LLMs are noisy sources, their output does contain false statements, yet they add more bits of signal than of noise relative to a human alone.

                  Or at least, they can add more add more bits of signal than of noise relative to a human alone, but humans who blindly copy-paste the output of an LLM without checking are a pain and add zero value to whatever situation they happen to be in.

                  For some hypothetical scenario, writing software because I know they can do that, asking an LLM to write some code for you may easily give you 10 kilobits of positive information (code that mostly works), and -30 bits of noise (each bit being one binary decision's worth of incorrect choice by the LLM in what to write, i.e. bugs); if you as a user don't know how to handle the -30 noise that could easily be a totally useless app, but if you can filter out 30 bits of noise, either manually because those 30 bits happen to be your skill set, or even in some cases by prompting it again with the failure mode, then you get to benefit from the 10 kilobits of good stuff that you didn't have before.

                  In many (but not all) cases, LLMs can fix more than 1 bit of mistakes per follow-up prompt.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.