‹ BackHN Continuity

Thread

Mercury 2.5 LLM hits 770 tokens per second

151 points · 92 comments · Retro_Dev

  1. bearjaws · · focus · HN ↗
    If you care about speed Cerebras gpt-oss-120b is 1400tk/s and "just as smart" in ranking.

    I've used it on a few for fun projects and its decent but the speed is crazy to watch.

    1. conception · · focus · HN ↗
      It was better when they had gemma at 1k. Inco does DS flash at about 600. A few places will do K3 and GLM in the hundreds.

      Such a tiny model at that t/s is less impressive than it would have been four months ago.

      1. physicallyIllfr · · focus · HN ↗
        Lighting my codebase on fire at the speed of light. Like microwaving the spaghetti.

        I genuinly only see these speeds being useful for customer service/transactional workflows. Of which much smaller models can do the job (but those dont make tons of money for companies like Cerebras that need to pay off massive amounts of debt).

        Nobody needs to code at 600 words per second. Using a 100tps model for an hour or so will leave you with 4-8hrs of code review and revision work.

        1. calgoo · · focus · HN ↗
          No you dont need to code at 600 w/s BUT at those speeds, you can start doing things like asking multiple different agents the same question and picking the best solution each time without noticing the lag.
          1. physicallyIllfr · · focus · HN ↗
            So 3x the code review lol. This is similar in concept to how a slot machine leta you choose 1x, 3x 6x lol
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.