‹ BackHN Continuity

Thread

Mercury 2.5 LLM hits 770 tokens per second

151 points · 92 comments · Retro_Dev

  1. bearjaws · · focus · HN ↗
    If you care about speed Cerebras gpt-oss-120b is 1400tk/s and "just as smart" in ranking.

    I've used it on a few for fun projects and its decent but the speed is crazy to watch.

    1. scosman · · focus · HN ↗
      Or better: Qwen 2.8 27b
      1. RussianCow · · focus · HN ↗
        Unfortunately, the lack of an input cache discount makes it prohibitively expensive for most use cases that aren't one-shot prompts.
        1. scosman · · focus · HN ↗
          well same applies to GPT OSS 120. Qwen is just the much smarter model of the 2 public options on Cerebras.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.