‹ BackHN Continuity

Thread

M5 Ultra Mac Studio Review

269 points · 262 comments · piotrgrabowski

  1. simonw · · focus · HN ↗
    The numbers I was most interested in are tucked away in a chart towards the bottom - the speed comparison of the Mac Studios v.s. a RTX 5090:

      Qwen3.8 27B tokens/sec generation speed
    
      Prompt size    8K    64K   128K   256K
      RTX 5090 PC    59    51    44     n/a
      M5 Ultra       48    39    32     24
      M3 Ultra       31    23.5  20     15
    
    A whole bunch more comparison numbers in this section: <a href="https:&#x2F;&#x2F;www.macstories.net&#x2F;stories&#x2F;m5-ultra-mac-studio-review-the-dream-mac-for-local-ai-agents&#x2F;#mx-pc" rel="nofollow">https:&#x2F;&#x2F;www.macstories.net&#x2F;stories&#x2F;m5-ultra-mac-studio-revie...
    1. redox99 · · focus · HN ↗
      A dense 27B doesn&#x27;t really make sense for the Mac. A MoE makes way more sense when you have modest bandwidth but lots of memory.
      1. peri-cl · · focus · HN ↗
        They do MoE. They benchmarked GLM 5.3-flash (320B &#x2F; 18B), and Qwen 3.8-flash-next (125B &#x2F; 6B). The dense Qwen is only focused (I assume) because it&#x27;s about the only thing that fits on a 5090, that they can compare the two heads on.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.