If you care about speed Cerebras gpt-oss-120b is 1400tk/s and "just as smart" in ranking.I've used it on a few for fun projects and its decent but the speed is crazy to watch.
Unfortunately, the lack of an input cache discount makes it prohibitively expensive for most use cases that aren't one-shot prompts.
well same applies to GPT OSS 120. Qwen is just the much smarter model of the 2 public options on Cerebras.
bearjaws · · focus · HN ↗
I've used it on a few for fun projects and its decent but the speed is crazy to watch.
scosman · · focus · HN ↗
RussianCow · · focus · HN ↗
scosman · · focus · HN ↗