‹ BackHN Continuity

Thread

Mercury 2.5 LLM hits 770 tokens per second

151 points · 92 comments · Retro_Dev

  1. walrus01 · · focus · HN ↗
    Pricing at $0.25 and $0.75 already puts its cost well above reasonably reputable inference providers for deepseek v4 flash or qwen 3.8-flash-next or similar class of open weight LLMs that fit in under 170GB of RAM, so I don't see the point. I think this is probably also stupider than laguna s 2.1 which can also be very cheap to serve.
    1. [deleted] · · focus · HN ↗

      [deleted]

    2. RussianCow · · focus · HN ↗
      The point is the speed.
      1. _aavaa_ · · focus · HN ↗
        If the model provides me with bad results because it's dumb, I don't care how quickly it does it.
        1. hermannj314 · · focus · HN ↗
          Is it possible to construct a control system where bad, fast and cheap can become good, fast, and cheap through repeated sampling and a strong spec/eval harness?

          I am trying to keep an open mind with AI, but I also have little understanding of control theory, trying to learn.

          1. freakynit · · focus · HN ↗
            Smaller models seems to get stuck in "loops" when you try to "handle" them this way.
          2. calgoo · · focus · HN ↗
            You can, but you need to break the problem into much smaller tasks, then check those answers, and finally have a harness that handles all the context, task breakup, task definitions, and validations each round.
        2. RussianCow · · focus · HN ↗
          But there are lots of use cases where a relatively "dumb" model is good enough.
        3. jgalt212 · · focus · HN ↗
          fast results that you need to verify are better than slow (allegedly better) results that you still need to verify. REPL vs batch.
          1. RussianCow · · focus · HN ↗
            Depends on how many times you need to iterate to get the result you want. If you need to run the fast model 5 times to get the results you need, compared to 1-2 times for a smarter but slower model, you've just eroded any advantage that the speed gave you.

            So, as always: it depends on the use case.

            1. jgalt212 · · focus · HN ↗
              fair enough, but in my experience fast models seem to be 5-10X faster than slow models.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.