The margins on NVidia datacenter hardware are ... high. At least one order of magnitude larger than a consumer chip.
Given the recent deepseekv4.1 advances - how good of a 3B model can we make to run on an iphone natively? is it good enough to match common muse/dot use cases for consumers? the phone is already always on.. no need for a cloud server.
Most phones are not really "always on" in any real sense, a phone on active standby uses very little power and most of it is for its mobile connection. Local AI is best run in a stationary homelab environment, even running it on laptops has its very real problems.
There are a few edge models, Spark-X2.5-4B, LFM2.5- in 8b-a1b and 2.6b variants, that are useful for on-device agent-y stuff, but I am doubtful that’s a valuable category. How many iPhone users have never opened Shortcuts in their life? Automating on-phone stuff seems niche.
I don’t expect the economics of local vs cloud ai to change until either the bubble pops or new ai chips land that can run big models fast with low power demand.
I think the interesting question is whether the average consumer cares? a 4B model should hit roughly the same numbers as GPT5.5 sometime in December/March. That is more than capable of performing a variety of complex work tasks.
I find this too. In fact, recently I've been pushing more and more to the latest and greatest model every time there is an update. It just saves me so much headache.
Right now you are right -- even if I ran my local LLM all day, the quality is not nearly as great, and it runs slowly -- so I use the tier one AI subscription services as they are faster and smarter. But that might only be true for a limited amount of time, and a limited number of circumstances.
To borrow your steam engine analogy, if local LLMs get as good as a Toyota Prius, even if OpenAI / Anthropic offer Ferraris, most people will be happy with their Prius as their daily driver.
Similarly, if the big labs start raising prices or cutting usage, you won't be able to use it as much as you want -- whereas a local LLM will run all day every day without costing you any extra money.
So right now you are right, but who knows how long that will last.
> For nearly all tasks, I want the fastest and smartest AI model.
This is not nearly true for everyone else in the world.
For example, think about the world in ~2021 pre-LLM. Would anyone say the sentence "I only want the fastest and smartest humans working on my project"?
No of course not. Most people don't want to pay $10 million dollar salary to the best programmers in the world. They prefer to pay $200k salary to a median programmer and that's good enough for their ecommerce website.
I think a scooter or a bus is a more apt comparison than a steam engine. The scooter and bus can both get the job done with some acceptable trade-offs, depending on your circumstances and what you’re willing to accept.
However, there are many use cases where they aren’t the right tool for the job.
This will change when out of control billing gets noticed. Then us devs will, I presume, get token rations.
I can see my new Thursday afternoon "oh chit" moment being that I didn't complete my weekly task because i torched all of those tokens M-W doing task/ticket grooming using the hot hot model instead of the dodo model with jira mcp connector :D
I would argue performance optimizations help with local/open models but hurt openai and anthropic - because open and or cheap/alternative models are threat to those companies. There is a fundamental contradiction/conflict between the prevelance of open models and the financial success of openai and anthropic. That is why they are doing everything they can to kill any open/cheap/efficient/chinese models (take a look at this thread - it was top of HN 40 mins ago, with very high engagement... now it is buried in page 5... totally normal and legit).
eggbrain · · focus · HN ↗
If local LLMs get "good" enough, people will soon paying for subscriptions to ChatGPT and Claude, which hurts their revenue.
ericol · · focus · HN ↗
Think you missed a word there.
kennywinker · · focus · HN ↗
lumost · · focus · HN ↗
Given the recent deepseekv4.1 advances - how good of a 3B model can we make to run on an iphone natively? is it good enough to match common muse/dot use cases for consumers? the phone is already always on.. no need for a cloud server.
christkv · · focus · HN ↗
zozbot234 · · focus · HN ↗
kennywinker · · focus · HN ↗
I don’t expect the economics of local vs cloud ai to change until either the bubble pops or new ai chips land that can run big models fast with low power demand.
lumost · · focus · HN ↗
kennywinker · · focus · HN ↗
You just can’t compress 1.5T into 4B without losing useful stuff.
But that doesn’t matter that much. If 4B can compose tools, retrieve info, and not get stuck in loops that is going to handle a lot of use cases
londons_explore · · focus · HN ↗
It is vanishingly rare I ask an older model to do any task. Newer bigger and smarter models will just do the task better.
Therefore, I believe we are nowhere near 'good enough'.
I never drive my steam engine to work these days. It isn't good enough.
danielmarkbruce · · focus · HN ↗
NoDodgeQuestion · · focus · HN ↗
eggbrain · · focus · HN ↗
To borrow your steam engine analogy, if local LLMs get as good as a Toyota Prius, even if OpenAI / Anthropic offer Ferraris, most people will be happy with their Prius as their daily driver.
Similarly, if the big labs start raising prices or cutting usage, you won't be able to use it as much as you want -- whereas a local LLM will run all day every day without costing you any extra money.
So right now you are right, but who knows how long that will last.
gretch · · focus · HN ↗
This is not nearly true for everyone else in the world.
For example, think about the world in ~2021 pre-LLM. Would anyone say the sentence "I only want the fastest and smartest humans working on my project"?
No of course not. Most people don't want to pay $10 million dollar salary to the best programmers in the world. They prefer to pay $200k salary to a median programmer and that's good enough for their ecommerce website.
tyre · · focus · HN ↗
But so many teams said they wanted to Raise the Bar to infinity and hire a World Class Team.
settsu · · focus · HN ↗
monkpit · · focus · HN ↗
However, there are many use cases where they aren’t the right tool for the job.
schmookeeg · · focus · HN ↗
I can see my new Thursday afternoon "oh chit" moment being that I didn't complete my weekly task because i torched all of those tokens M-W doing task/ticket grooming using the hot hot model instead of the dodo model with jira mcp connector :D
aleqs · · focus · HN ↗