‹ BackHN Continuity

Thread

Mercury 2.5 LLM hits 770 tokens per second

151 points · 92 comments · Retro_Dev

  1. bearjaws · · focus · HN ↗
    If you care about speed Cerebras gpt-oss-120b is 1400tk/s and "just as smart" in ranking.

    I've used it on a few for fun projects and its decent but the speed is crazy to watch.

    1. ford · · focus · HN ↗
      Also kimi 2.6 at 1000tps (as of may), though when we reached out they had a >12 month waitlist and minimum 7-8 figure annual token spend.

      [0] <a href="https:&#x2F;&#x2F;www.cerebras.ai&#x2F;blog&#x2F;cerebras-kimi-k2-Enterprise" rel="nofollow">https:&#x2F;&#x2F;www.cerebras.ai&#x2F;blog&#x2F;cerebras-kimi-k2-Enterprise

      1. walrus01 · · focus · HN ↗
        7-8 figures annual spend will buy a hell of a lot of capable local inference hardware you can own, though it won&#x27;t be at the absurd token&#x2F;s rate, you&#x27;ll be able to run almost anything on it... And it&#x27;ll still have a good residual resale value after 4 years the way things are going now.
        1. podocarp · · focus · HN ↗
          Feel like you could spend 6 figures building out a team and the rest renting compute for a whole year, and get the team to create a local inference solution with that kind of budget…
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.