The margins on NVidia datacenter hardware are ... high. At least one order of magnitude larger than a consumer chip.
Given the recent deepseekv4.1 advances - how good of a 3B model can we make to run on an iphone natively? is it good enough to match common muse/dot use cases for consumers? the phone is already always on.. no need for a cloud server.
There are a few edge models, Spark-X2.5-4B, LFM2.5- in 8b-a1b and 2.6b variants, that are useful for on-device agent-y stuff, but I am doubtful that’s a valuable category. How many iPhone users have never opened Shortcuts in their life? Automating on-phone stuff seems niche.
I don’t expect the economics of local vs cloud ai to change until either the bubble pops or new ai chips land that can run big models fast with low power demand.
I think the interesting question is whether the average consumer cares? a 4B model should hit roughly the same numbers as GPT5.5 sometime in December/March. That is more than capable of performing a variety of complex work tasks.
eggbrain · · focus · HN ↗
If local LLMs get "good" enough, people will soon paying for subscriptions to ChatGPT and Claude, which hurts their revenue.
kennywinker · · focus · HN ↗
lumost · · focus · HN ↗
Given the recent deepseekv4.1 advances - how good of a 3B model can we make to run on an iphone natively? is it good enough to match common muse/dot use cases for consumers? the phone is already always on.. no need for a cloud server.
kennywinker · · focus · HN ↗
I don’t expect the economics of local vs cloud ai to change until either the bubble pops or new ai chips land that can run big models fast with low power demand.
lumost · · focus · HN ↗
kennywinker · · focus · HN ↗
You just can’t compress 1.5T into 4B without losing useful stuff.
But that doesn’t matter that much. If 4B can compose tools, retrieve info, and not get stuck in loops that is going to handle a lot of use cases