‹ BackHN Continuity

Thread

M5 Ultra Mac Studio Review

269 points · 262 comments · piotrgrabowski

  1. simonw · · focus · HN ↗
    The numbers I was most interested in are tucked away in a chart towards the bottom - the speed comparison of the Mac Studios v.s. a RTX 5090:

      Qwen3.8 27B tokens/sec generation speed
    
      Prompt size    8K    64K   128K   256K
      RTX 5090 PC    59    51    44     n/a
      M5 Ultra       48    39    32     24
      M3 Ultra       31    23.5  20     15
    
    A whole bunch more comparison numbers in this section: <a href="https:&#x2F;&#x2F;www.macstories.net&#x2F;stories&#x2F;m5-ultra-mac-studio-review-the-dream-mac-for-local-ai-agents&#x2F;#mx-pc" rel="nofollow">https:&#x2F;&#x2F;www.macstories.net&#x2F;stories&#x2F;m5-ultra-mac-studio-revie...
    1. nacs · · focus · HN ↗
      That&#x27;s a dense model. Of course it will do worse.

      Now try running that Qwen 3.8 Next model on the 5090 and tell me what TPS you get (hint: it&#x27;s near 0 since it doesnt fit the 32GB VRAM on 5090 vs the 256 in OPs M5).

      1. peri-cl · · focus · HN ↗
        Surprisingly, the Reddit crowd are reporting 50–60 tokens&#x2F;s (for the 32 GiB 5090 + 128 GiB RAM)—on par with the M5 Ultra benchmarks, despite both the PCIe bottleneck and much smaller DDR5 bandwidth,

        <a href="https:&#x2F;&#x2F;old.reddit.com&#x2F;r&#x2F;LocalLLaMA&#x2F;comments&#x2F;1wl06np&#x2F;qwen38flashnext_on_1x_rtx_5090_tg50_ts_pp2300_ts&#x2F;" rel="nofollow">https:&#x2F;&#x2F;old.reddit.com&#x2F;r&#x2F;LocalLLaMA&#x2F;comments&#x2F;1wl06np&#x2F;qwen38f...

        (Note it&#x27;s a sparse MoE with only 6B active).

        1. bitexploder · · focus · HN ↗
          I have a 3 year old gaming system. RTX 4080 w&#x2F;128GB of DDR5. It runs Qwen 38 Flash around 44-40 t&#x2F;s with 128K context. It is on a specialized build that caches MoE experts and uses an optimized 3bit quant that basically is within a few points of the full 8 bit quant. In general, in casual benchmarking with Alibaba&#x27;s endpoint I could not tell much of a difference. Overall this model is very good on long horizon agentic work. The main pain point for it is that its input processing speed is slow. Regardless, it gets meaningful work done.

          I paid $500 for the RAM in Nov 2023 :)

          1. peri-cl · · focus · HN ↗
            &gt; &quot;I paid $500 for the RAM in Nov 2023 :)&quot;

            No wonder Warren Buffet gave up and resigned.

            1. bitexploder · · focus · HN ↗
              Right? To get comparable brand and quality DDR5, which isn’t particularly great at AI anything it is ~$2000. All you had to do was start hoarding 3090 GPU and RAM in 2023. It is unhinged.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.