‹ BackHN Continuity

Thread

M5 Ultra Mac Studio Review

269 points · 262 comments · piotrgrabowski

  1. simonw · · focus · HN ↗
    The numbers I was most interested in are tucked away in a chart towards the bottom - the speed comparison of the Mac Studios v.s. a RTX 5090:

      Qwen3.8 27B tokens/sec generation speed
    
      Prompt size    8K    64K   128K   256K
      RTX 5090 PC    59    51    44     n/a
      M5 Ultra       48    39    32     24
      M3 Ultra       31    23.5  20     15
    
    A whole bunch more comparison numbers in this section: <a href="https:&#x2F;&#x2F;www.macstories.net&#x2F;stories&#x2F;m5-ultra-mac-studio-review-the-dream-mac-for-local-ai-agents&#x2F;#mx-pc" rel="nofollow">https:&#x2F;&#x2F;www.macstories.net&#x2F;stories&#x2F;m5-ultra-mac-studio-revie...
    1. alex7o · · focus · HN ↗
      On my m5 max 27b model does 75tps on 256k ctx and starts at 80 on the 8k ctx when you add <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;collections&#x2F;z-lab&#x2F;dflash-2" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;collections&#x2F;z-lab&#x2F;dflash-2 to it. So yeah base might be 30tps (I used iq4) but mtp or dflash help a lot and should be used when checking what is useful and what is not for running models as it is not fare to judge without them.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.