The author observes that a call to GPT-5.6 Luna is only 4-5 orders of magnitude more expensive than grep, and then predicts that at current rates of progress, calling an LLM will soon be cheaper than a grep. I think this is a good time to invoke Stein's Law: "If something cannot go on forever, it will stop." These efficiency improvements won't continue forever. It's more likely that the per-call cost of high-quality, compiled software like grep will be a lower-bound that LLMs asymptotically approach, rather than a line that they blow past with perpetual exponential progress. (Barring a true breakthrough in something like quantum computing or room-temperature superconductors.)
Yeah I bumped on that too. If it's possible to make llms cheaper than current grep, then it is also almost certainly possible to make grep cheaper.
You can burn anything* into an ASIC to make it cheaper per-call.
non-backreferencing grep is not very difficult to implement in an ASIC either. But it's probably not worth it because of how relatively rarely you use it and of the data transfer costs.
LLMs are great candidates for ASIC-burning because they're slow compared even to network speeds and run all the time. The issue is that you don't want to burn a specific model or architecture that then becomes obsolete.
So you've got two possible futures, and both guarantee large price drops: (a) LLMs keep getting better and better and better, so ability/$ keeps rising; or (b) LLMs plateau in ability, in which they will start getting ASIC'd.
This whole story really reminds me of crypto coins. Like.. going from mining one coin, or lets say token, to millions of fractions like 0.00000000001 bitcoin a week.
I probably replied to the wrong comment with my previous response. The low lying fruit, in this case, should be going from single cell to multi cell, but that took ages.
It accelerated from there rather than slowing down.
And the progress humanity made in the last 100 years alone... some things end up to be completely world-changing because they enable things that only were an unfeasible dream before. The invention of the printing press, sanitation, vaccines , computers, the Internet certainly are such enabler technologies.
With AI, the question is still open if this will actually turn out to be something useful or if it will in the end just be another way for the elites to make untold profits.
> LLMs plateau in ability, in which they will start getting ASIC'd.
They don't need to plateu for that to happen. There are companies already building AI on ASIC, and IIRC they were approach 12 months lead time. A 12 months old frontier model (Sonnet 4.5, GPT-5, Kimi K2) for 1% of the price is still a rather good value proposition.
It will be a better value proposition now than it was 12 months ago. It's likely to be better yet in another 12 months. There may be room for a parallel to Moore's Law here.
Right now, I do actually use OpenAI's gpt-oss-safeguard-20b for somethings, was released 11 months ago, and is $0.075/M input / $0.30/M output now. I could see this model being in fairly widespread use at 10x speed and 1/10th cost if it was introduced today. Meaning, that for some usecases (moderation) i think dedicated chips can pan out today.
But for more general models, its tougher. Gemini 3 pro was launched in November, if ASICs brought it down 1/10th in cost, it would be $0.20/$1.2. GPT 6 Luna is $0.1/$0.50. Luna is better at a lot of things, but not everything. So 1/10th doesn't really make the ASICS investment worth it in my opinion, but if it brought it down to 1% ($0.02 / $0.12) it would be a really compelling model with a lot of use.
BUT, do i think something like Luna is probably generally capable of doing a huge amount of knowledge work. So if Luna came out at 1/10th the cost a year from now, it would probably be compelling for a while.
It all depends on the rate of improvement in cost/capability.
In order for the ASIC to achieve 1% of the price, there'd have to be 99% overhead in GPU implementations which for some reason you'd have to be able to eliminate in ASICs but not in GPUs. That seems rather implausible.
To be clear, I agree with the overall premise of the article!
But I would probably take a long horizon bet that the grep implementation on my machine will remain cheaper than an equivalent ai task, even though I think those ai tasks will become far cheaper over time.
I just think the original comment's model of asymptotic approach is probably more likely to be accurate than the model of the line blowing through this grep-like cost level.
It's almost a certainty that LLMs have a "core" that will essentially never become obsolete, possibly even 80% to 90% of their parameters. The rules of English and other languages, core ideas in math and science, all of history, nearly all literature, etc. We don't really understand what's going on inside LLMs enough yet to make good use of this, but one day we will have "core logic" neural networks with stable weights burned into ASIC that are doing the heavy lifting, with more dynamic continually-tuned models manipulating the inputs and outputs into those core models. There are also likely stable expert models on topics that don't change much that we could already do this with.
Inference costs cannot keep falling forever, but they do still have a long way to go.
No, it just gets much more complicated to implement grep both in general, and especially in hardware if you support backreferences, since those make it impossible to compile the regular expression into a state machine.
Grep (or ripgrep at least) is i/o bottlenecked at this point. It's impossible to process data at faster than i/o speeds, since you have to get the data to the processor somehow. That doesn't change whether that processing is grep on a CPU, or LLM on an ASIC.
i/o could still be sped up though? And I dunno, maybe this actually is an argument for how the llm version could end up being faster, because there is lots of investment in crazy fast i/o hardware and protocols to get the data into the chips.
I did a toy project once implementing a limited version of grep on an FPGA and was able to get some speedup over GNU grep at the time, though marginal.
And a patch panel and a bunch of patch cables that you plug in and out to construct your pipelines.
And eventually hire people whose job it is to patch pipelines on demand for everyone in the office.
“Hey Jim, I’m gonna output the systemd logs of nginx on line five, can you assemble a grep pipeline for me to match all HTTP 500 status codes from /api/cart POST request log lines? Connect the filtered output to Tim’s desk, line 7. He’s there now, we are trying to figure something out.”
It was a companion of punchcards, the tabulating machine. Programming by wires.
<a href="https://en.wikipedia.org/wiki/Tabulating_machine#Selected_models_and_timeline" rel="nofollow">https://en.wikipedia.org/wiki/Tabulating_machine#Selected_mo...
Almost 30 years ago there was a project about making a computer which used FPGAs for all "software". It was called RAW.
Baring it all to software: Raw machines | IEEE Journals & Magazine | IEEE Xplore <a href="https://share.google/nI9GyFJu4HvbYFrBf" rel="nofollow">https://share.google/nI9GyFJu4HvbYFrBf
(If you Google the name you will find free PDFs as well, the IEEE page is more useful as a summary and such.)
I wonder how close you can get with Nvidia's GPUDirect. Hook the fast NVME directly up to the GPU (well, it gets direct DMA to GPU at any rate), then implement parallel grep in CUDA... profit?
GPUs and inference ASICS also have large amounts of high bandwidth memory, plus lots of high speed storage cache, and dedicated very high bandwidth scale-out and scale-up networks. Because they are also often bound by I/O bandwidth.
If your problem is grepping crazy amounts of data, the infrastructure for LLMs isn't a bad place to look for an example.
I wouldn't be surprised if that's what Apple is focused on for their next generation platforms – I wonder if more layers of caching between their SSDs and unified memory are on the cards.
depends on what you are grepping ... greapping a large file might be more expensive one day than generating n-th token with LLM that works fully in hardware
you could make hardware implementation of grep and store the file itself next to it in some ROM but that's not a very useful grep ... while hardware LLM is exactly as useful as software LLM only orders of magnitude faster
jetrink · · focus · HN ↗
The author observes that a call to GPT-5.6 Luna is only 4-5 orders of magnitude more expensive than grep, and then predicts that at current rates of progress, calling an LLM will soon be cheaper than a grep. I think this is a good time to invoke Stein's Law: "If something cannot go on forever, it will stop." These efficiency improvements won't continue forever. It's more likely that the per-call cost of high-quality, compiled software like grep will be a lower-bound that LLMs asymptotically approach, rather than a line that they blow past with perpetual exponential progress. (Barring a true breakthrough in something like quantum computing or room-temperature superconductors.)
sanderjd · · focus · HN ↗
serbuvlad · · focus · HN ↗
non-backreferencing grep is not very difficult to implement in an ASIC either. But it's probably not worth it because of how relatively rarely you use it and of the data transfer costs.
LLMs are great candidates for ASIC-burning because they're slow compared even to network speeds and run all the time. The issue is that you don't want to burn a specific model or architecture that then becomes obsolete.
So you've got two possible futures, and both guarantee large price drops: (a) LLMs keep getting better and better and better, so ability/$ keeps rising; or (b) LLMs plateau in ability, in which they will start getting ASIC'd.
thenthenthen · · focus · HN ↗
pixl97 · · focus · HN ↗
hobofan · · focus · HN ↗
Yes, because it was a largely random undirected process.
pixl97 · · focus · HN ↗
Then you have to climb to another branch to get more fruit. The biggest issue with most problem space discovery is you're doing it blindfolded.
onraglanroad · · focus · HN ↗
It accelerated from there rather than slowing down.
onraglanroad · · focus · HN ↗
The complexity gets faster as you get on with it.
mschuster91 · · focus · HN ↗
With AI, the question is still open if this will actually turn out to be something useful or if it will in the end just be another way for the elites to make untold profits.
formerly_proven · · focus · HN ↗
hobofan · · focus · HN ↗
They don't need to plateu for that to happen. There are companies already building AI on ASIC, and IIRC they were approach 12 months lead time. A 12 months old frontier model (Sonnet 4.5, GPT-5, Kimi K2) for 1% of the price is still a rather good value proposition.
cestith · · focus · HN ↗
mchusma · · focus · HN ↗
Right now, I do actually use OpenAI's gpt-oss-safeguard-20b for somethings, was released 11 months ago, and is $0.075/M input / $0.30/M output now. I could see this model being in fairly widespread use at 10x speed and 1/10th cost if it was introduced today. Meaning, that for some usecases (moderation) i think dedicated chips can pan out today.
But for more general models, its tougher. Gemini 3 pro was launched in November, if ASICs brought it down 1/10th in cost, it would be $0.20/$1.2. GPT 6 Luna is $0.1/$0.50. Luna is better at a lot of things, but not everything. So 1/10th doesn't really make the ASICS investment worth it in my opinion, but if it brought it down to 1% ($0.02 / $0.12) it would be a really compelling model with a lot of use.
BUT, do i think something like Luna is probably generally capable of doing a huge amount of knowledge work. So if Luna came out at 1/10th the cost a year from now, it would probably be compelling for a while.
It all depends on the rate of improvement in cost/capability.
sanderjd · · focus · HN ↗
atq2119 · · focus · HN ↗
sanderjd · · focus · HN ↗
But I would probably take a long horizon bet that the grep implementation on my machine will remain cheaper than an equivalent ai task, even though I think those ai tasks will become far cheaper over time.
I just think the original comment's model of asymptotic approach is probably more likely to be accurate than the model of the line blowing through this grep-like cost level.
feoren · · focus · HN ↗
Inference costs cannot keep falling forever, but they do still have a long way to go.
vrighter · · focus · HN ↗
lstodd · · focus · HN ↗
serbuvlad · · focus · HN ↗
nvme0n1p1 · · focus · HN ↗
sanderjd · · focus · HN ↗
serbuvlad · · focus · HN ↗
I did a toy project once implementing a limited version of grep on an FPGA and was able to get some speedup over GNU grep at the time, though marginal.
In any case, LLMs aren't IO bound :))
sanderjd · · focus · HN ↗
cobbal · · focus · HN ↗
bee_rider · · focus · HN ↗
benterix · · focus · HN ↗
QuantumNomad_ · · focus · HN ↗
And eventually hire people whose job it is to patch pipelines on demand for everyone in the office.
“Hey Jim, I’m gonna output the systemd logs of nginx on line five, can you assemble a grep pipeline for me to match all HTTP 500 status codes from /api/cart POST request log lines? Connect the filtered output to Tim’s desk, line 7. He’s there now, we are trying to figure something out.”
“Sure thing Bob, give me a moment.”
kgwgk · · focus · HN ↗
bartread · · focus · HN ↗
formerly_proven · · focus · HN ↗
nxobject · · focus · HN ↗
serbuvlad · · focus · HN ↗
justsid · · focus · HN ↗
xigoi · · focus · HN ↗
mkesper · · focus · HN ↗
Retro_Dev · · focus · HN ↗
Scarblac · · focus · HN ↗
mpweiher · · focus · HN ↗
Put a Transputer in a Lego brick. Not "like a Lego brick", an actual Lego brick. Turn the notches into connectors for the serial link.
Plug'n'play.
sanderjd · · focus · HN ↗
klipt · · focus · HN ↗
lou1306 · · focus · HN ↗
mhast · · focus · HN ↗
Baring it all to software: Raw machines | IEEE Journals & Magazine | IEEE Xplore <a href="https://share.google/nI9GyFJu4HvbYFrBf" rel="nofollow">https://share.google/nI9GyFJu4HvbYFrBf
(If you Google the name you will find free PDFs as well, the IEEE page is more useful as a summary and such.)
nxobject · · focus · HN ↗
NooneAtAll3 · · focus · HN ↗
swiftcoder · · focus · HN ↗
alexshendi · · focus · HN ↗
dgacmu · · focus · HN ↗
podocarp · · focus · HN ↗
readams · · focus · HN ↗
If your problem is grepping crazy amounts of data, the infrastructure for LLMs isn't a bad place to look for an example.
nxobject · · focus · HN ↗
I wouldn't be surprised if that's what Apple is focused on for their next generation platforms – I wonder if more layers of caching between their SSDs and unified memory are on the cards.
BiteCode_dev · · focus · HN ↗
scotty79 · · focus · HN ↗
you could make hardware implementation of grep and store the file itself next to it in some ROM but that's not a very useful grep ... while hardware LLM is exactly as useful as software LLM only orders of magnitude faster
sanderjd · · focus · HN ↗
rdsubhas · · focus · HN ↗
When the author wrote Llm can be as cheap as a tool, I read it as not equivalent. They even said the Llm can be embedded into a tool.
Their point was, the higher level use case — like classification — could become as cheap as grep. Which is quite well possible.
sanderjd · · focus · HN ↗