What's really fun about this is that when people hear "open models," they rarely question where the inference happens, and who owns it, and how that can feed RL.
I am a huge Anthropic fan, USA fan, but are we cooked with AI? Sorry SI, that's the important thing.
The open model excitement isn't directly regarding cloud services.
The direct excitement regarding open models is that many of these Flash Mixture-of-Expert models run reasonably well on hardware a tech employee in the West, and businesses in less affluent countries, can realistically afford.
The indirect excitement is that the models are so efficient, so cloud prices also end up being very low.
You don't see the same scale of excitement surrounding the open weight trillion+ parameter models because, while it's neat they're open weight, it doesn't mean a lot if you need $50k worth of computers to just barely run them.
> You don't see the same scale of excitement surrounding the open weight trillion+ parameter models because, while it's neat they're open weight, it doesn't mean a lot if you need $50k worth of computers to just barely run them.
This looks like a rather outdated POV. With increasingly pervasive use of SSD offload, there nothing particularly stopping you from running even the largest open models on ordinary local hardware. Sure, it will be really slow, but if you need the smarts for e.g. a one-off planning role it's a no-brainer.
I think SSD offload will become more feasible an the engram approach matures.
For now though, I would question if the difference in smarts is large enough to justify tolerating a generation speed measured in seconds per token compared to spending more turns refining a plan with a flash model.
consumer451 · · focus · HN ↗
I am a huge Anthropic fan, USA fan, but are we cooked with AI? Sorry SI, that's the important thing.
hgoel · · focus · HN ↗
The direct excitement regarding open models is that many of these Flash Mixture-of-Expert models run reasonably well on hardware a tech employee in the West, and businesses in less affluent countries, can realistically afford.
The indirect excitement is that the models are so efficient, so cloud prices also end up being very low.
You don't see the same scale of excitement surrounding the open weight trillion+ parameter models because, while it's neat they're open weight, it doesn't mean a lot if you need $50k worth of computers to just barely run them.
zozbot234 · · focus · HN ↗
This looks like a rather outdated POV. With increasingly pervasive use of SSD offload, there nothing particularly stopping you from running even the largest open models on ordinary local hardware. Sure, it will be really slow, but if you need the smarts for e.g. a one-off planning role it's a no-brainer.
hgoel · · focus · HN ↗
For now though, I would question if the difference in smarts is large enough to justify tolerating a generation speed measured in seconds per token compared to spending more turns refining a plan with a flash model.