A whole bunch more comparison numbers in this section: <a href="https://www.macstories.net/stories/m5-ultra-mac-studio-review-the-dream-mac-for-local-ai-agents/#mx-pc" rel="nofollow">https://www.macstories.net/stories/m5-ultra-mac-studio-revie...
I dont' know why people spend huge money on these and Spark. The 5090 is running qwen 3.8 at 200+ tps!! That's 1-2 orders of magnitude faster.
This 1000%. Data centres don't equate to medium sized labs and businesses. A stack of Macs is up and running without digging trenches, an electrician on staff and a department of PhDs to justify the spend.
It's likely that a stack of Macs will draw more power for slower prefill/decode than equivalently priced Nvidia GPUs. If power efficient inference is the goal, Macs are a non-starter.
Oh, it's absolutely both. The power you waste waiting for TFTT on prefill will absolutely compound at the "medium sized labs and businesses" scale.
simonw · · focus · HN ↗
gpugreg · · focus · HN ↗
beastman82 · · focus · HN ↗
I dont' know why people spend huge money on these and Spark. The 5090 is running qwen 3.8 at 200+ tps!! That's 1-2 orders of magnitude faster.
mathisfun123 · · focus · HN ↗
throwaway27448 · · focus · HN ↗
_hugerobots_ · · focus · HN ↗
bigyabai · · focus · HN ↗
_hugerobots_ · · focus · HN ↗
bigyabai · · focus · HN ↗