‹ BackHN Continuity

Thread

Mercury 2.5 LLM hits 770 tokens per second

151 points · 92 comments · Retro_Dev

  1. nylonstrung · · focus · HN ↗
    I honestly think the diffusion LLM approach is a dead end

    It's telling that frontier labs like Google toyed around with it but didn't invest further even for their most speed and cost sensitive small models

    Still unclear for what, if any use cases this is pareto frontier

    1. vineyardmike · · focus · HN ↗
      From what I’ve heard, the issue is more that it’s harder to efficiently share the hardware across diffusion requests, so it’s more expensive to serve.

      It sounds like there might be opportunities for local models (not open weight, but actually locally run) to use diffusion for faster responses on weaker hardware that doesn’t need to be shared.

      But yea, it’s still a red-ish flag that big labs haven’t invested much in it. I could see Google/Apple getting value of this sort of local model, but maybe there’s enough research behind traditional models that it’s not worth the distraction at this point in time.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.