Taalas (Acquired by AMD, back in August) created Jimmy[0], a little chat app that runs on a POC chip with ~14k tps. Yes, 14,000 tokens per second. Sure, it's just a 8B model or so (Llama 3.1 8B), but I can imagine that having a 1.58-bit model might be helpful for their next chip.
Heck, what would happen if you used a dLLM (d for diffusion)?
infogulch · · focus · HN ↗
If ternary llms work out and are baked into hardware as custom silicon I bet they'll be shockingly efficient.
NostraDavid · · focus · HN ↗
Taalas (Acquired by AMD, back in August) created Jimmy[0], a little chat app that runs on a POC chip with ~14k tps. Yes, 14,000 tokens per second. Sure, it's just a 8B model or so (Llama 3.1 8B), but I can imagine that having a 1.58-bit model might be helpful for their next chip.
Heck, what would happen if you used a dLLM (d for diffusion)?
[0]: <a href="https://chatjimmy.ai/" rel="nofollow">https://chatjimmy.ai/
nwah1 · · focus · HN ↗