‹ BackHN Continuity

Thread

M5 Ultra Mac Studio Review

269 points · 262 comments · piotrgrabowski

  1. simonw · · focus · HN ↗
    The numbers I was most interested in are tucked away in a chart towards the bottom - the speed comparison of the Mac Studios v.s. a RTX 5090:

      Qwen3.8 27B tokens/sec generation speed
    
      Prompt size    8K    64K   128K   256K
      RTX 5090 PC    59    51    44     n/a
      M5 Ultra       48    39    32     24
      M3 Ultra       31    23.5  20     15
    
    A whole bunch more comparison numbers in this section: <a href="https:&#x2F;&#x2F;www.macstories.net&#x2F;stories&#x2F;m5-ultra-mac-studio-review-the-dream-mac-for-local-ai-agents&#x2F;#mx-pc" rel="nofollow">https:&#x2F;&#x2F;www.macstories.net&#x2F;stories&#x2F;m5-ultra-mac-studio-revie...
    1. nacs · · focus · HN ↗
      That&#x27;s a dense model. Of course it will do worse.

      Now try running that Qwen 3.8 Next model on the 5090 and tell me what TPS you get (hint: it&#x27;s near 0 since it doesnt fit the 32GB VRAM on 5090 vs the 256 in OPs M5).

      1. peri-cl · · focus · HN ↗
        Surprisingly, the Reddit crowd are reporting 50–60 tokens&#x2F;s (for the 32 GiB 5090 + 128 GiB RAM)—on par with the M5 Ultra benchmarks, despite both the PCIe bottleneck and much smaller DDR5 bandwidth,

        <a href="https:&#x2F;&#x2F;old.reddit.com&#x2F;r&#x2F;LocalLLaMA&#x2F;comments&#x2F;1wl06np&#x2F;qwen38flashnext_on_1x_rtx_5090_tg50_ts_pp2300_ts&#x2F;" rel="nofollow">https:&#x2F;&#x2F;old.reddit.com&#x2F;r&#x2F;LocalLLaMA&#x2F;comments&#x2F;1wl06np&#x2F;qwen38f...

        (Note it&#x27;s a sparse MoE with only 6B active).

        1. nacs · · focus · HN ↗
          Good to know thanks.

          That&#x27;s with CPU offload to a DDR5 6000 RAM though which is around $3-4k at least.

          1. well_ackshually · · focus · HN ↗
            Unlike a 256GB M5 Ultra that is $10k+.
            1. nacs · · focus · HN ↗
              Apple product won&#x27;t be the cheapest but it is a full package (CPU, RAM, VRAM&#x2F;GPU, fast-storage, etc).

              If you look at the pricing of a full (x86) AI workstation you&#x27;d need around the nvidia GPU, you&#x27;d approach $10k easily (and be using a ton more wattage too).

            2. cma · · focus · HN ↗
              But the 5090 they use there is now going low stock and selling for over $7500 in some places.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.