I'm still sad that we haven't seen a new Taalas style chip a la <a href="https://chatjimmy.ai/" rel="nofollow">https://chatjimmy.ai/. Smaller models are good enough now to make that insane burst of tokens so useful.
I don't know the model behind this, but it is absurdly bad.
> Write me a coherent paragraph in French, without ever using the letter "e".
> Voilà une phrase claire et concise : "Le village est situé dans les montagnes. Le soleil est haut. Il y a des animaux dans le village. Il pleut dans les montagnes."
I suppose this is just a demo of how fast an LLM can be, I wonder if there are tradeoffs with larger/smarter models. Also, for a human usage, at what point are tokens generated fast enough that it's pretty much instant? My bet is below 1000 tps
To be fair, you picked a well-known tricky benchmark for LLMs: When working on an embedding spelling disappears after the embedding level. I imagine modern frontier models have tools that let them read back their input to work around this issue.
jjcm · · focus · HN ↗
pil0u · · focus · HN ↗
> Write me a coherent paragraph in French, without ever using the letter "e".
> Voilà une phrase claire et concise : "Le village est situé dans les montagnes. Le soleil est haut. Il y a des animaux dans le village. Il pleut dans les montagnes."
I suppose this is just a demo of how fast an LLM can be, I wonder if there are tradeoffs with larger/smarter models. Also, for a human usage, at what point are tokens generated fast enough that it's pretty much instant? My bet is below 1000 tps
fph · · focus · HN ↗