What's really fun about this is that when people hear "open models," they rarely question where the inference happens, and who owns it, and how that can feed RL.
I am a huge Anthropic fan, USA fan, but are we cooked with AI? Sorry SI, that's the important thing.
The open model excitement isn't directly regarding cloud services.
The direct excitement regarding open models is that many of these Flash Mixture-of-Expert models run reasonably well on hardware a tech employee in the West, and businesses in less affluent countries, can realistically afford.
The indirect excitement is that the models are so efficient, so cloud prices also end up being very low.
You don't see the same scale of excitement surrounding the open weight trillion+ parameter models because, while it's neat they're open weight, it doesn't mean a lot if you need $50k worth of computers to just barely run them.
I believe that I might agree with you. That is the attraction, and the truth, and the marketing. Near frontier open models are truly amazing.
Avoiding 2 labs having control of knowledge work at the true frontier is not just good for nerd reasons. If only 2 labs control the frontier of knowledge work, well.. that's pretty much the end of the USA's market/service economy. I mean if one or two labs run it, does that make a country's economy?
Meanwhile, we are a deeply stupid species. Based on multiple previous conversations on this website, it appears that for example, z.ai hosting is not a big deal.
Yes, anyone doing real due diligence on client data will face reality, maybe. However, given the long tail of external consultancies, the CCP is going to eat it all due to our outsourced laziness. Shareholder value, amirite!!?
This is how we lost our manufacturing base. Why would the token manufacturing base be any different? As far as I can tell, we have gotten even stupider in the last few years. I guess changing the name to SI and tariffs on Canada and the EU will solve these problems.
I think most posts are best read when there is a solution at the end, but is there one? What would that look like?
I want to add to this: we didn't lose our manufacturing base just for shareholder value. The consumers also got super cheap and yet exponentially more amazing technology.
I am an old-ish person. My first PC cost as much as a used car. Thanks to global supply chains, we all get those in our pockets, all the time! Capitalism at its best!
If we could agree to stop killing each other, there is literally no end to this hockey stick.
This is the true contrarian POV in 2026, fight me.
What I mean is, Thiel/Yarvin followers. Let's have a talk about this.
> You don't see the same scale of excitement surrounding the open weight trillion+ parameter models because, while it's neat they're open weight, it doesn't mean a lot if you need $50k worth of computers to just barely run them.
This looks like a rather outdated POV. With increasingly pervasive use of SSD offload, there nothing particularly stopping you from running even the largest open models on ordinary local hardware. Sure, it will be really slow, but if you need the smarts for e.g. a one-off planning role it's a no-brainer.
Yes, but there is only one reason that one might buy truly private inference. That is true privacy.
The price one pays for this is that lower than frontier models are dumber. At the current pace of advancement, is that a good idea? Isn't it better to just go with with ZDR contracts, and not lose in the software product game, which now moves at relativistic speed?
I think SSD offload will become more feasible an the engram approach matures.
For now though, I would question if the difference in smarts is large enough to justify tolerating a generation speed measured in seconds per token compared to spending more turns refining a plan with a flash model.
consumer451 · · focus · HN ↗
I am a huge Anthropic fan, USA fan, but are we cooked with AI? Sorry SI, that's the important thing.
hgoel · · focus · HN ↗
The direct excitement regarding open models is that many of these Flash Mixture-of-Expert models run reasonably well on hardware a tech employee in the West, and businesses in less affluent countries, can realistically afford.
The indirect excitement is that the models are so efficient, so cloud prices also end up being very low.
You don't see the same scale of excitement surrounding the open weight trillion+ parameter models because, while it's neat they're open weight, it doesn't mean a lot if you need $50k worth of computers to just barely run them.
consumer451 · · focus · HN ↗
Avoiding 2 labs having control of knowledge work at the true frontier is not just good for nerd reasons. If only 2 labs control the frontier of knowledge work, well.. that's pretty much the end of the USA's market/service economy. I mean if one or two labs run it, does that make a country's economy?
Meanwhile, we are a deeply stupid species. Based on multiple previous conversations on this website, it appears that for example, z.ai hosting is not a big deal.
Yes, anyone doing real due diligence on client data will face reality, maybe. However, given the long tail of external consultancies, the CCP is going to eat it all due to our outsourced laziness. Shareholder value, amirite!!?
This is how we lost our manufacturing base. Why would the token manufacturing base be any different? As far as I can tell, we have gotten even stupider in the last few years. I guess changing the name to SI and tariffs on Canada and the EU will solve these problems.
I think most posts are best read when there is a solution at the end, but is there one? What would that look like?
consumer451 · · focus · HN ↗
I am an old-ish person. My first PC cost as much as a used car. Thanks to global supply chains, we all get those in our pockets, all the time! Capitalism at its best!
If we could agree to stop killing each other, there is literally no end to this hockey stick.
This is the true contrarian POV in 2026, fight me.
What I mean is, Thiel/Yarvin followers. Let's have a talk about this.
[deleted] · · focus · HN ↗
[deleted]
zozbot234 · · focus · HN ↗
This looks like a rather outdated POV. With increasingly pervasive use of SSD offload, there nothing particularly stopping you from running even the largest open models on ordinary local hardware. Sure, it will be really slow, but if you need the smarts for e.g. a one-off planning role it's a no-brainer.
consumer451 · · focus · HN ↗
The price one pays for this is that lower than frontier models are dumber. At the current pace of advancement, is that a good idea? Isn't it better to just go with with ZDR contracts, and not lose in the software product game, which now moves at relativistic speed?
hgoel · · focus · HN ↗
For now though, I would question if the difference in smarts is large enough to justify tolerating a generation speed measured in seconds per token compared to spending more turns refining a plan with a flash model.