‹ BackHN Continuity

Thread

Mercury 2.5 LLM hits 770 tokens per second

151 points · 92 comments · Retro_Dev

  1. rvz · · focus · HN ↗
    The speed means absolutely nothing when it is finishing almost dead last when compared to the frontier AI companies.
    1. copperx · · focus · HN ↗
      Ah, the old "good, fast, or cheap; pick two" proves true once again.
      1. downrightmike · · focus · HN ↗
        Give it a few months.
    2. glouwbug · · focus · HN ↗
      Some of us want fast food
    3. voiceeh · · focus · HN ↗
      Not if your use case needs speed. For one of my products I can't use an LLM that has a p99 of >700ms for TTFT.
      1. hansvm · · focus · HN ↗
        If it could output 1k tokens per second but needed 4 seconds to produce the first batch of 4k, would that not be viable?
    4. timClicks · · focus · HN ↗
      It means something, because it an iterative workflow. If you're willing to burn tokens, it's possible for weaker models to implement tasks by incrementally improving drafts.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.