‹ BackHN Continuity

Thread

Mercury 2.5 LLM hits 770 tokens per second

151 points · 92 comments · Retro_Dev

  1. bearjaws · · focus · HN ↗
    If you care about speed Cerebras gpt-oss-120b is 1400tk/s and "just as smart" in ranking.

    I've used it on a few for fun projects and its decent but the speed is crazy to watch.

    1. conception · · focus · HN ↗
      It was better when they had gemma at 1k. Inco does DS flash at about 600. A few places will do K3 and GLM in the hundreds.

      Such a tiny model at that t/s is less impressive than it would have been four months ago.

      1. physicallyIllfr · · focus · HN ↗
        Lighting my codebase on fire at the speed of light. Like microwaving the spaghetti.

        I genuinly only see these speeds being useful for customer service/transactional workflows. Of which much smaller models can do the job (but those dont make tons of money for companies like Cerebras that need to pay off massive amounts of debt).

        Nobody needs to code at 600 words per second. Using a 100tps model for an hour or so will leave you with 4-8hrs of code review and revision work.

        1. RugnirViking · · focus · HN ↗
          > Nobody needs to code at 600 words per second.

          I do. I used to use haiku for the speed. Now its just as slow as the rest. Speed is my #1 ranking of how good a model is

          1. physicallyIllfr · · focus · HN ↗
            I prefer a fast model too, but you cannot get more done just because its faster. You just get to the human parts a bit faster. Code review, revision ect.
            1. RugnirViking · · focus · HN ↗
              absolutely not ! it's not an "ohhh ill get 100x more coding done" it's definitely a preference/work style thing. I find it's hard to get in "the flow" when managing multiple agents, and if I just do one at a time with today's frontier models, I find myself waiting 10 minutes twiddling my thumbs while they do something all the time. Then needing to catch myself and turn back over to it when its done, all these micro switches between tasks is really difficult for me. If I had an instant agent I probably would be somewhat faster, but not zomg1000xunicornrockstar nonsense. I would be way happier though
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.