A whole bunch more comparison numbers in this section: <a href="https://www.macstories.net/stories/m5-ultra-mac-studio-review-the-dream-mac-for-local-ai-agents/#mx-pc" rel="nofollow">https://www.macstories.net/stories/m5-ultra-mac-studio-revie...
I assume those are non-batched. I think the M series GPU can do 4X to 8X depending on model quant, which means if you can batch queries you'll get almost 4X to 8X performance.
simonw · · focus · HN ↗
api · · focus · HN ↗