‹ BackHN Continuity

Thread

M5 Ultra Mac Studio Review

269 points · 262 comments · piotrgrabowski

  1. simonw · · focus · HN ↗
    The numbers I was most interested in are tucked away in a chart towards the bottom - the speed comparison of the Mac Studios v.s. a RTX 5090:

      Qwen3.8 27B tokens/sec generation speed
    
      Prompt size    8K    64K   128K   256K
      RTX 5090 PC    59    51    44     n/a
      M5 Ultra       48    39    32     24
      M3 Ultra       31    23.5  20     15
    
    A whole bunch more comparison numbers in this section: <a href="https:&#x2F;&#x2F;www.macstories.net&#x2F;stories&#x2F;m5-ultra-mac-studio-review-the-dream-mac-for-local-ai-agents&#x2F;#mx-pc" rel="nofollow">https:&#x2F;&#x2F;www.macstories.net&#x2F;stories&#x2F;m5-ultra-mac-studio-revie...
    1. gpugreg · · focus · HN ↗
      Those RTX 5090 numbers are bad. You can get over 200 tps with ninfer using NVFP4 and MTP.
      1. beastman82 · · focus · HN ↗
        can confirm.

        I dont&#x27; know why people spend huge money on these and Spark. The 5090 is running qwen 3.8 at 200+ tps!! That&#x27;s 1-2 orders of magnitude faster.

        1. tomega2134 · · focus · HN ↗
          Is a 5090 still cost efficent when it is (currently) unobtainable? Or when obtainable only at current prices (min. $6500 USD)?
          1. throwaway219450 · · focus · HN ↗
            Personally I think the price is way too high right now. It’s a power hungry gaming GPU. The efficient single card equivalent would be a 4500 Blackwell which launched at about $3500. Or you could get a 9700 32GB or an Arc B70 for well under $2k, today. You only buy a 5090 if you want absolute speed.

            32GB is still not that much. I would rather get a Spark and have the RAM to experiment with larger LLMs, even if it was slow.

            1. searealist · · focus · HN ↗
              A 5090 has 2x tensor cores and 2x bandwidth and can be run at 400W (2x watts).
              1. throwaway219450 · · focus · HN ↗
                Being fast and having a power target doesn’t mean it’s cost efficient though. I would pay the launch cost for one, but not 3-4x inflated.
                1. searealist · · focus · HN ↗
                  How does this relate to 4500 vs 5090? I&#x27;m just pointing out that 5090 likely has twice the performance of the 4500 and likely maintains that at 2x watts if you want.
                  1. throwaway219450 · · focus · HN ↗
                    You didn&#x27;t specify in your earlier post, so I wasn&#x27;t sure exactly which comparison you were making. But yeah, the perf&#x2F;watt actually looks the same for those, so the cost per token evens out. It is nice not having to manage 400+W though. I like the 4000 for that reason, it&#x27;s effectively a 3090 that runs at half the TDP.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.