A whole bunch more comparison numbers in this section: <a href="https://www.macstories.net/stories/m5-ultra-mac-studio-review-the-dream-mac-for-local-ai-agents/#mx-pc" rel="nofollow">https://www.macstories.net/stories/m5-ultra-mac-studio-revie...
I dont' know why people spend huge money on these and Spark. The 5090 is running qwen 3.8 at 200+ tps!! That's 1-2 orders of magnitude faster.
A 5090 has a 1.79TB/s memory bandwidth. Qwen 3.8 27B NVFP4 is 22GB. You cannot generate tokens faster than the weights can traverse the GPU memory, so that makes max generation speed without MTP to be 81T/s. Say MTP is giving you 0.5 acceptance rate (very good), that is 1.5 * 81 is 121T/s. Even with a perfect acceptance rate you would only get 162T/s.
simonw · · focus · HN ↗
gpugreg · · focus · HN ↗
beastman82 · · focus · HN ↗
I dont' know why people spend huge money on these and Spark. The 5090 is running qwen 3.8 at 200+ tps!! That's 1-2 orders of magnitude faster.
Eisenstein · · focus · HN ↗
medvezhenok · · focus · HN ↗
spider-mario · · focus · HN ↗