‹ BackHN Continuity

Thread

Mercury 2.5 LLM hits 770 tokens per second

151 points · 92 comments · Retro_Dev

  1. bearjaws · · focus · HN ↗
    If you care about speed Cerebras gpt-oss-120b is 1400tk/s and "just as smart" in ranking.

    I've used it on a few for fun projects and its decent but the speed is crazy to watch.

    1. ford · · focus · HN ↗
      Also kimi 2.6 at 1000tps (as of may), though when we reached out they had a >12 month waitlist and minimum 7-8 figure annual token spend.

      [0] <a href="https:&#x2F;&#x2F;www.cerebras.ai&#x2F;blog&#x2F;cerebras-kimi-k2-Enterprise" rel="nofollow">https:&#x2F;&#x2F;www.cerebras.ai&#x2F;blog&#x2F;cerebras-kimi-k2-Enterprise

      1. sharktheone · · focus · HN ↗
        yeah. K2.6 can run on insane speeds. So sad that they don&#x27;t have K3 yet.

        But it can apparently also run 5.6 Sol

        1. eli · · focus · HN ↗
          Yeah but at “call to discuss pricing” rates
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.