‹ BackHN Continuity

Thread

Mercury 2.5 LLM hits 770 tokens per second

151 points · 92 comments · Retro_Dev

  1. bearjaws · · focus · HN ↗
    If you care about speed Cerebras gpt-oss-120b is 1400tk/s and "just as smart" in ranking.

    I've used it on a few for fun projects and its decent but the speed is crazy to watch.

    1. LoganDark · · focus · HN ↗
      Please do not try to use gpt-oss-120b over Cerebras. It is broken, screws up tool calls most of the time, forgets to end thinking blocks and has all sorts of other issues. The speed is amazing but it is absolutely not worth it, especially at that quite incredible cost. Think: $5–10/minute levels of cost with a single agent, because Cerebras also offers no cache pricing for input tokens at all.
      1. eli · · focus · HN ↗
        Which is wild because it does, in fact, do caching
        1. fakwandi_priv · · focus · HN ↗
          > There is no additional fee for using prompt caching. Input tokens, whether served from the cache or processed fresh, are billed at the standard input token rate for the respective model.

          So what he’s saying is correct, there is no separate cache pricing, which by normal standards should be 10% of the cost, which can become exceedingly expensive for anything other than single turn. The way they are stating this is of course strange..

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.