‹ BackHN Continuity

Thread

M5 Ultra Mac Studio Review

269 points · 262 comments · piotrgrabowski

  1. simonw · · focus · HN ↗
    The numbers I was most interested in are tucked away in a chart towards the bottom - the speed comparison of the Mac Studios v.s. a RTX 5090:

      Qwen3.8 27B tokens/sec generation speed
    
      Prompt size    8K    64K   128K   256K
      RTX 5090 PC    59    51    44     n/a
      M5 Ultra       48    39    32     24
      M3 Ultra       31    23.5  20     15
    
    A whole bunch more comparison numbers in this section: <a href="https:&#x2F;&#x2F;www.macstories.net&#x2F;stories&#x2F;m5-ultra-mac-studio-review-the-dream-mac-for-local-ai-agents&#x2F;#mx-pc" rel="nofollow">https:&#x2F;&#x2F;www.macstories.net&#x2F;stories&#x2F;m5-ultra-mac-studio-revie...
    1. gpugreg · · focus · HN ↗
      Those RTX 5090 numbers are bad. You can get over 200 tps with ninfer using NVFP4 and MTP.
      1. beastman82 · · focus · HN ↗
        can confirm.

        I dont&#x27; know why people spend huge money on these and Spark. The 5090 is running qwen 3.8 at 200+ tps!! That&#x27;s 1-2 orders of magnitude faster.

        1. fhub · · focus · HN ↗
          How are you deciding which work to send to the 5090 vs a frontier model, or making the two work together nicely?

          Correct is much more important than fast for me, but if I could get correct and fast, that would obviously be amazing.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.