A whole bunch more comparison numbers in this section: <a href="https://www.macstories.net/stories/m5-ultra-mac-studio-review-the-dream-mac-for-local-ai-agents/#mx-pc" rel="nofollow">https://www.macstories.net/stories/m5-ultra-mac-studio-revie...
I dont' know why people spend huge money on these and Spark. The 5090 is running qwen 3.8 at 200+ tps!! That's 1-2 orders of magnitude faster.
Exactly; when I first got my RTX 5070 Ti (16gb, to game with!!!, upgrading from VEGA56), I loaded then-latest Qwen3.6 (~30B, cannot remember exactly). My only prior LLM experience was with models <8gb, primarily llama3.1.
My technical-expert twin played around with these LLMs, for about an hour, and then correctly reasoned "it's able to be WRONG, faster."
This seems apt. My next LLM machine will be closer to 96gb+ vRAM.
simonw · · focus · HN ↗
gpugreg · · focus · HN ↗
beastman82 · · focus · HN ↗
I dont' know why people spend huge money on these and Spark. The 5090 is running qwen 3.8 at 200+ tps!! That's 1-2 orders of magnitude faster.
nacs · · focus · HN ↗
ProllyInfamous · · focus · HN ↗
My technical-expert twin played around with these LLMs, for about an hour, and then correctly reasoned "it's able to be WRONG, faster."
This seems apt. My next LLM machine will be closer to 96gb+ vRAM.
selectodude · · focus · HN ↗
khriss · · focus · HN ↗
ProllyInfamous · · focus · HN ↗
...actually: get help.