If you care about speed Cerebras gpt-oss-120b is 1400tk/s and "just as smart" in ranking.I've used it on a few for fun projects and its decent but the speed is crazy to watch.
It was better when they had gemma at 1k. Inco does DS flash at about 600. A few places will do K3 and GLM in the hundreds.Such a tiny model at that t/s is less impressive than it would have been four months ago.
Inco sucks. I tried their GLM 5.3 Flash and it was quantized to the point of hallucinating Chinese in the middle of English only agentic sessions. Never happened with any other provider.
bearjaws · · focus · HN ↗
I've used it on a few for fun projects and its decent but the speed is crazy to watch.
conception · · focus · HN ↗
Such a tiny model at that t/s is less impressive than it would have been four months ago.
lostmsu · · focus · HN ↗