‹ BackHN Continuity

Thread

Mercury 2.5 LLM hits 770 tokens per second

151 points · 92 comments · Retro_Dev

  1. bearjaws · · focus · HN ↗
    If you care about speed Cerebras gpt-oss-120b is 1400tk/s and "just as smart" in ranking.

    I've used it on a few for fun projects and its decent but the speed is crazy to watch.

    1. LoganDark · · focus · HN ↗
      Please do not try to use gpt-oss-120b over Cerebras. It is broken, screws up tool calls most of the time, forgets to end thinking blocks and has all sorts of other issues. The speed is amazing but it is absolutely not worth it, especially at that quite incredible cost. Think: $5–10/minute levels of cost with a single agent, because Cerebras also offers no cache pricing for input tokens at all.
      1. bearjaws · · focus · HN ↗
        Not been my experience, I have it using tool calls in a video game I am building and it correctly adheres ~99% of the time.

        I have it retry on failure, but you should do that with any LLM really.

        1. LoganDark · · focus · HN ↗
          I kept having experiences with gpt-oss-120b on Cerebras where it would get stuck in a thinking block and then start endlessly saying things like "Running the command now." or "Making the changes now." and then simply repeating similar sentences like that forever instead of actually making the tool call. It made tool calls other times, so it wasn't an issue with tool calls being impossible, but it just wasn't doing a good job of using them for real instead of simply saying it would. So this was not an issue of it starting a tool call and then putting invalid syntax inside of it, it just would not make the tool call it was supposed to whatsoever. There's no automatic way to retry that.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.