‹ BackHN Continuity

Thread

Tokens too cheap to meter

354 points · 227 comments · teoruiz

  1. jetrink · · focus · HN ↗
    > Tokens become cheaper than tool calls

    The author observes that a call to GPT-5.6 Luna is only 4-5 orders of magnitude more expensive than grep, and then predicts that at current rates of progress, calling an LLM will soon be cheaper than a grep. I think this is a good time to invoke Stein's Law: "If something cannot go on forever, it will stop." These efficiency improvements won't continue forever. It's more likely that the per-call cost of high-quality, compiled software like grep will be a lower-bound that LLMs asymptotically approach, rather than a line that they blow past with perpetual exponential progress. (Barring a true breakthrough in something like quantum computing or room-temperature superconductors.)

    1. sanderjd · · focus · HN ↗
      Yeah I bumped on that too. If it's possible to make llms cheaper than current grep, then it is also almost certainly possible to make grep cheaper.
      1. serbuvlad · · focus · HN ↗
        You can burn anything* into an ASIC to make it cheaper per-call.

        non-backreferencing grep is not very difficult to implement in an ASIC either. But it's probably not worth it because of how relatively rarely you use it and of the data transfer costs.

        LLMs are great candidates for ASIC-burning because they're slow compared even to network speeds and run all the time. The issue is that you don't want to burn a specific model or architecture that then becomes obsolete.

        So you've got two possible futures, and both guarantee large price drops: (a) LLMs keep getting better and better and better, so ability/$ keeps rising; or (b) LLMs plateau in ability, in which they will start getting ASIC'd.

        1. nvme0n1p1 · · focus · HN ↗
          Grep (or ripgrep at least) is i/o bottlenecked at this point. It's impossible to process data at faster than i/o speeds, since you have to get the data to the processor somehow. That doesn't change whether that processing is grep on a CPU, or LLM on an ASIC.
          1. readams · · focus · HN ↗
            GPUs and inference ASICS also have large amounts of high bandwidth memory, plus lots of high speed storage cache, and dedicated very high bandwidth scale-out and scale-up networks. Because they are also often bound by I/O bandwidth.

            If your problem is grepping crazy amounts of data, the infrastructure for LLMs isn't a bad place to look for an example.

            1. nxobject · · focus · HN ↗
              > plus lots of high speed storage cache

              I wouldn't be surprised if that's what Apple is focused on for their next generation platforms – I wonder if more layers of caching between their SSDs and unified memory are on the cards.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.