‹ BackHN Continuity

Thread

M5 Ultra Mac Studio Review

269 points · 262 comments · piotrgrabowski

  1. simonw · · focus · HN ↗
    The numbers I was most interested in are tucked away in a chart towards the bottom - the speed comparison of the Mac Studios v.s. a RTX 5090:

      Qwen3.8 27B tokens/sec generation speed
    
      Prompt size    8K    64K   128K   256K
      RTX 5090 PC    59    51    44     n/a
      M5 Ultra       48    39    32     24
      M3 Ultra       31    23.5  20     15
    
    A whole bunch more comparison numbers in this section: <a href="https:&#x2F;&#x2F;www.macstories.net&#x2F;stories&#x2F;m5-ultra-mac-studio-review-the-dream-mac-for-local-ai-agents&#x2F;#mx-pc" rel="nofollow">https:&#x2F;&#x2F;www.macstories.net&#x2F;stories&#x2F;m5-ultra-mac-studio-revie...
    1. lhl · · focus · HN ↗
      A basic llama-bench on Qwen 3.8 27B UD-Q4_K_M gives pp512 3920 tok&#x2F;s &#x2F; tg128 81 tok&#x2F;s on a 500W RTX PRO 6000 (should be similar speeds to a 5090, chip is basically the same, just less VRAM). With MTP3 this is 140 tok&#x2F;s on mtp-bench.

      This is with llama.cpp. You can of course use vLLM&#x2F;SGLang well on these cards and they&#x27;re even faster. On vLLM w&#x2F; NVIDIA&#x2F;Qwen3.8-27B-NVFP4 baseline has a prefill of about 13,000 tok&#x2F;s. The baseline tok&#x2F;s is 72 tok&#x2F;s, but at mtp7, it&#x27;s 157 tok&#x2F;s, and w&#x2F; dflash7 that goes up to 215 tok&#x2F;s. On mtp-bench, DFlash2 gets a hair under 300 tok&#x2F;s w&#x2F; the code_python prompt.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.