‹ BackHN Continuity

Thread

M5 Ultra Mac Studio Review

269 points · 262 comments · piotrgrabowski

  1. simonw · · focus · HN ↗
    The numbers I was most interested in are tucked away in a chart towards the bottom - the speed comparison of the Mac Studios v.s. a RTX 5090:

      Qwen3.8 27B tokens/sec generation speed
    
      Prompt size    8K    64K   128K   256K
      RTX 5090 PC    59    51    44     n/a
      M5 Ultra       48    39    32     24
      M3 Ultra       31    23.5  20     15
    
    A whole bunch more comparison numbers in this section: <a href="https:&#x2F;&#x2F;www.macstories.net&#x2F;stories&#x2F;m5-ultra-mac-studio-review-the-dream-mac-for-local-ai-agents&#x2F;#mx-pc" rel="nofollow">https:&#x2F;&#x2F;www.macstories.net&#x2F;stories&#x2F;m5-ultra-mac-studio-revie...
    1. gpugreg · · focus · HN ↗
      Those RTX 5090 numbers are bad. You can get over 200 tps with ninfer using NVFP4 and MTP.
      1. ActorNightly · · focus · HN ↗

        [dead]

        1. abletonlive · · focus · HN ↗
          Tok&#x2F;sec is 0 on a 3090 for most of the models that the mac can run
          1. ActorNightly · · focus · HN ↗

            [dead]

            1. abletonlive · · focus · HN ↗
              &gt; So given that, which one of these is true about you?

              Well, if those are the only two options you can come up with it&#x27;s pretty clear that this isn&#x27;t about me or what I am, you have a false model of reality.

              &gt; Running very large models on Mac is unusable at 10 tok&#x2F;sec.

              There are plenty of examples of models running at well over 10 tok&#x2F;sec that aren&#x27;t viable on the 3090. In fact such examples are found in the review in the OP. Did you not read the article?

              I think you&#x27;re projecting pretty hard with the two options you&#x27;ve listed. Go touch some grass, you seem overly frustrated that reality doesn&#x27;t meet your expectations.

              1. ActorNightly · · focus · HN ↗
                Since you clearly don&#x27;t use local llms, allow me to educate you - anything under 100 tok&#x2F;sec is USELESS. When you are coding, the idea is that you want to have a system that can generate files fast, hopefully correct on the first try. Cloud models do this. Local models, by nature of having less parameters and more quantization, often require more guidance and repeated inference to get it right. The antigenic harnesses that people set up around local llms leverage this.

                Looking at the article, which you clearly didn&#x27;t read,the m5 ultra runs Qwen3.8, which fits on one GPU conveniently, at ~20 tok&#x2F;sec. This is a fucking joke. It will take roughly a minute to generate one code file. Congrats if you want privacy I guess, but for straight up coding, you are better just using cloud models.

                Meanwhile, I have an $800 mini PC, $200 Occulink gpu dock, a $2000 3090 and a $300 power supply, and I can run Qwen at over 100 tok&#x2F;sec prefill, not to mention insanely quicker during inference. So its pointless to spend Mac M5 Ultra prices on Apple shit when they can have something much faster for cheaper

                The whole thing of &quot;well I can run bigger models that don&#x27;t fit on a GPU&quot; is either paid Apple advertising, or you are just an igorant fanboy.

                So I ask you again, which one are you?

                1. abletonlive · · focus · HN ↗
                  &gt; allow me to educate you

                  No thanks, you&#x27;re not in a position to do that clearly.

                  &gt; Since you clearly don&#x27;t use local llms

                  I do, probably a lot longer than you have actually.

                  &gt; anything under 100 tok&#x2F;sec is USELESS

                  Objectively wrong. You sound like you&#x27;re really behind and you&#x27;re so myopic that you think coding is the only use case for local LLMs. I&#x27;m a professional software dev and that&#x27;s the least interesting use case of local LLMs.

                  &gt; Looking at the article, which you clearly didn&#x27;t read,the m5 ultra runs Qwen3.8, which fits on one GPU conveniently, at ~20 tok&#x2F;sec.

                  You clearly didn&#x27;t read the article or have reading comprehension issues. The model is Qwen3.8-Flash-Next 4 and 5-bit quant, neither of which &quot;conveniently fits on one GPU&quot;. Sorry that your hardware doesn&#x27;t live up to your own delusions and can&#x27;t even run Qwen3.8-Flash-Next at 4&#x2F;5 bit quant. You are taking the Quen3.8-27B numbers, something that the article isn&#x27;t really that concerned with, and trying to make it fit into your narrative.

                  &gt; So I ask you again, which one are you?

                  Well I&#x27;m someone that suggests that you should touch some grass and reevaluate your personal issues. You seem angry. Perhaps it&#x27;s best to figure your own issues before trying to figure out why people are excited about Apple hardware for local llms. I am sure the people that need to interact with you in society would be very grateful if you took the time to do this.

                  1. ActorNightly · · focus · HN ↗
                    Nice try.

                    A) He literally says &quot;I tested a different Qwen model for the comparisons between Mac and PC.&quot; The model he tested has to fit on one GPU, otherwise the inference is dogshit slow as you are offloading results to ram. If you ran any amount of local inference, you would know this. Considering that Qwen3.8-Flash-Next Q4 is still 100gb, there is no realistic way to run this with a 5090. The model that was run was this <a href="https:&#x2F;&#x2F;ollama.com&#x2F;library&#x2F;qwen3.8:27b">https:&#x2F;&#x2F;ollama.com&#x2F;library&#x2F;qwen3.8:27b. And the speed of that model on a 5090 in terms of tok&#x2F;sec is not 60 lol.

                    B) If M5 ultra runs 40 tok&#x2F;sec on qwen3.8:27b (and lets assume its the mlx version to gain a performance boost: <a href="https:&#x2F;&#x2F;ollama.com&#x2F;library&#x2F;qwen3.8:27b-mlx">https:&#x2F;&#x2F;ollama.com&#x2F;library&#x2F;qwen3.8:27b-mlx), you have to be delusional to believe it can run 100gb models at 100 tok&#x2F;sec lol.

                    As a bonus, in terms of use, its pretty well known that Qwen models are RLed to chase benchmarks. Check out <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;Qwen&#x2F;Qwen3.8-27B" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;Qwen&#x2F;Qwen3.8-27B versus <a href="https:&#x2F;&#x2F;qwen.ai&#x2F;blog?id=qwen3.8-flash-next" rel="nofollow">https:&#x2F;&#x2F;qwen.ai&#x2F;blog?id=qwen3.8-flash-next, using different benchmarks the 27b outperforms the flash next on agentic coding. But it matches it in other areas pretty well. So tell me again why you need 100gb models running dogshit slow at peak ~20 tok&#x2F;sec?

                    It is so incredibly sad how hard you try to sound intelligent. But thats on par for the course of any person hyping up apple products, throughout apples history.

                    Considering that Apple probably doesn&#x27;t want you to engage in this level of pettiness for their advertising posts, you have outed yourself to be #2. And Im not angry at all lol, you keep doing what you do, people like you in the industry are the reason I can work 8 hours a week and still get get paid a lot while being reviewed highly.

                    1. abletonlive · · focus · HN ↗
                      You&#x27;re straight up wrong and it&#x27;s hilarious. Seriously, go touch grass. I feel bad for the people that have to interact with you in real life because you must be a miserable person.

                      You are comparing dense models to MoE. You can&#x27;t just take a qhen3.8:27b, a dense model, and extrapolate that math to make your arguments for the MoE model being discussed in the review.

                      All this takes is a simple bit of research, it&#x27;s not like this author is the only person that&#x27;s ever used qwen3.8-flash-next or another MoE on a mac ultra. <a href="https:&#x2F;&#x2F;tbreak.com&#x2F;mac-studio-m5-max-review-local-ai&#x2F;" rel="nofollow">https:&#x2F;&#x2F;tbreak.com&#x2F;mac-studio-m5-max-review-local-ai&#x2F;

                      &gt; people like you in the industry are the reason I can work 8 hours a week and still get get paid a lot while being reviewed highly.

                      LOL. No offense but you sound like a carriage return. I&#x27;m pretty sure you sit at your computer hitting enter on a mechanical keyboard next to your fabulous mini pc living in 2025. You couldn&#x27;t even imagine another use case besides being a carriage return for a local llm.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.