‹ BackHN Continuity

Thread

M5 Ultra Mac Studio Review

269 points · 262 comments · piotrgrabowski

  1. simonw · · focus · HN ↗
    The numbers I was most interested in are tucked away in a chart towards the bottom - the speed comparison of the Mac Studios v.s. a RTX 5090:

      Qwen3.8 27B tokens/sec generation speed
    
      Prompt size    8K    64K   128K   256K
      RTX 5090 PC    59    51    44     n/a
      M5 Ultra       48    39    32     24
      M3 Ultra       31    23.5  20     15
    
    A whole bunch more comparison numbers in this section: <a href="https:&#x2F;&#x2F;www.macstories.net&#x2F;stories&#x2F;m5-ultra-mac-studio-review-the-dream-mac-for-local-ai-agents&#x2F;#mx-pc" rel="nofollow">https:&#x2F;&#x2F;www.macstories.net&#x2F;stories&#x2F;m5-ultra-mac-studio-revie...
    1. gpugreg · · focus · HN ↗
      Those RTX 5090 numbers are bad. You can get over 200 tps with ninfer using NVFP4 and MTP.
      1. beastman82 · · focus · HN ↗
        can confirm.

        I dont&#x27; know why people spend huge money on these and Spark. The 5090 is running qwen 3.8 at 200+ tps!! That&#x27;s 1-2 orders of magnitude faster.

        1. Eisenstein · · focus · HN ↗
          A 5090 has a 1.79TB&#x2F;s memory bandwidth. Qwen 3.8 27B NVFP4 is 22GB. You cannot generate tokens faster than the weights can traverse the GPU memory, so that makes max generation speed without MTP to be 81T&#x2F;s. Say MTP is giving you 0.5 acceptance rate (very good), that is 1.5 * 81 is 121T&#x2F;s. Even with a perfect acceptance rate you would only get 162T&#x2F;s.
          1. beastman82 · · focus · HN ↗
            Off the top of my head, I&#x27;m guessing we&#x27;re missing sparse attention. But I&#x27;ll run your challenge through and see where the gaps are. I promise I&#x27;m telling the truth :)
            1. [deleted] · · focus · HN ↗

              [deleted]

          2. medvezhenok · · focus · HN ↗
            I think you’re missing that MTP can predict more than 1 token in advance.
            1. spider-mario · · focus · HN ↗
              In fact, isn’t that the “M” in “MTP”?
          3. girvo · · focus · HN ↗
            It really does get it, because MTP is usually run at &quot;3 token&quot; depth. It&#x27;s pretty shocking to watch
            1. [deleted] · · focus · HN ↗

              [deleted]

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.