‹ BackHN Continuity

Thread

Mercury 2.5 LLM hits 770 tokens per second

151 points · 92 comments · Retro_Dev

  1. jjcm · · focus · HN ↗
    I&#x27;m still sad that we haven&#x27;t seen a new Taalas style chip a la <a href="https:&#x2F;&#x2F;chatjimmy.ai&#x2F;" rel="nofollow">https:&#x2F;&#x2F;chatjimmy.ai&#x2F;. Smaller models are good enough now to make that insane burst of tokens so useful.
    1. pil0u · · focus · HN ↗
      I don&#x27;t know the model behind this, but it is absurdly bad.

      &gt; Write me a coherent paragraph in French, without ever using the letter &quot;e&quot;.

      &gt; Voilà une phrase claire et concise : &quot;Le village est situé dans les montagnes. Le soleil est haut. Il y a des animaux dans le village. Il pleut dans les montagnes.&quot;

      I suppose this is just a demo of how fast an LLM can be, I wonder if there are tradeoffs with larger&#x2F;smarter models. Also, for a human usage, at what point are tokens generated fast enough that it&#x27;s pretty much instant? My bet is below 1000 tps

      1. fransje26 · · focus · HN ↗
        Then again, good luck writing a coherent paragraph in French without an &quot;e&quot;. :-)
        1. [deleted] · · focus · HN ↗

          [deleted]

        2. mejutoco · · focus · HN ↗
          There is a book written under this premise. Probably the inspiration for that prompt

          <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;A_Void" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;A_Void

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.