‹ BackHN Continuity

Thread

M5 Ultra Mac Studio Review

269 points · 262 comments · piotrgrabowski

  1. simonw · · focus · HN ↗
    The numbers I was most interested in are tucked away in a chart towards the bottom - the speed comparison of the Mac Studios v.s. a RTX 5090:

      Qwen3.8 27B tokens/sec generation speed
    
      Prompt size    8K    64K   128K   256K
      RTX 5090 PC    59    51    44     n/a
      M5 Ultra       48    39    32     24
      M3 Ultra       31    23.5  20     15
    
    A whole bunch more comparison numbers in this section: <a href="https:&#x2F;&#x2F;www.macstories.net&#x2F;stories&#x2F;m5-ultra-mac-studio-review-the-dream-mac-for-local-ai-agents&#x2F;#mx-pc" rel="nofollow">https:&#x2F;&#x2F;www.macstories.net&#x2F;stories&#x2F;m5-ultra-mac-studio-revie...
    1. redox99 · · focus · HN ↗
      A dense 27B doesn&#x27;t really make sense for the Mac. A MoE makes way more sense when you have modest bandwidth but lots of memory.
      1. tcdent · · focus · HN ↗
        A dense model (up to the amount of memory available) actually does make the most sense on unified memory architectures. But when you hit the limit of what you can hold in memory, you reach the limitation of the platform.

        Whereas a hybrid architecture with distinct DRAM and VRAM with sparse MoE, you can leverage two different bit rates depending on the actual need for constant access to common layers versus sparse access to infrequent layers and arbitrage the difference in cost for each of those in distinct classes of hardware.

        1. nojs · · focus · HN ↗
          &gt; A dense model (up to the amount of memory available) actually does make the most sense on unified memory architectures

          Inference time is going to be dominated by the low memory bandwidth on these Macs, so a dense model will suffer most. It’s more of an opportunity for large MoE models with a low number of active experts since you can keep all experts in VRAM but not pay the bandwidth cost until they are used.

          &gt; you can leverage two different bit rates depending on the actual need for constant access to common layers versus sparse access to infrequent layers

          This is an interesting direction that I expect to see more of. But for most models currently you need basically all experts loaded since they are chosen per token.

          Apple seems to be researching longer horizon expert caching, where they keep experts swapped in for longer runs of tokens [1]. Other labs are offloading ngram caches but not sure if they’re pursuing anything like this?

          1. <a href="https:&#x2F;&#x2F;machinelearning.apple.com&#x2F;research&#x2F;introducing-third-generation-of-apple-foundation-models" rel="nofollow">https:&#x2F;&#x2F;machinelearning.apple.com&#x2F;research&#x2F;introducing-third...

          1. thejazzman · · focus · HN ↗
            1.2TB&#x2F;s is already considered slow? Things are moving quickly!
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.