A whole bunch more comparison numbers in this section: <a href="https://www.macstories.net/stories/m5-ultra-mac-studio-review-the-dream-mac-for-local-ai-agents/#mx-pc" rel="nofollow">https://www.macstories.net/stories/m5-ultra-mac-studio-revie...
That's a dense model. Of course it will do worse.
Now try running that Qwen 3.8 Next model on the 5090 and tell me what TPS you get (hint: it's near 0 since it doesnt fit the 32GB VRAM on 5090 vs the 256 in OPs M5).
Surprisingly, the Reddit crowd are reporting 50–60 tokens/s (for the 32 GiB 5090 + 128 GiB RAM)—on par with the M5 Ultra benchmarks, despite both the PCIe bottleneck and much smaller DDR5 bandwidth,
simonw · · focus · HN ↗
nacs · · focus · HN ↗
Now try running that Qwen 3.8 Next model on the 5090 and tell me what TPS you get (hint: it's near 0 since it doesnt fit the 32GB VRAM on 5090 vs the 256 in OPs M5).
peri-cl · · focus · HN ↗
<a href="https://old.reddit.com/r/LocalLLaMA/comments/1wl06np/qwen38flashnext_on_1x_rtx_5090_tg50_ts_pp2300_ts/" rel="nofollow">https://old.reddit.com/r/LocalLLaMA/comments/1wl06np/qwen38f...
(Note it's a sparse MoE with only 6B active).
nacs · · focus · HN ↗
That's with CPU offload to a DDR5 6000 RAM though which is around $3-4k at least.
well_ackshually · · focus · HN ↗
cma · · focus · HN ↗