‹ BackHN Continuity

Thread

M5 Ultra Mac Studio Review

269 points · 262 comments · piotrgrabowski

  1. srcreigh · · focus · HN ↗
    This is great as a first look, but the author is not a developer, so we don't yet know whether a dev can be as productive with local models on M5 Mac Studio compared to a 20x subscription plan.

    I'm also curious about any new low hanging optimization opportunities in the kernels for this new hardware.

    It's already clear to me that M5 Mac Studio is more cost-effective than anything you can run on open router, assuming decent utilization.

    The M5 Mac Studio will be the most cost effective way to run uncensored cyber capable open agents.

    An exciting tipping point will be if programmers can get an Astra-Ultra like experience all week with this hardware. That would be a real sense where this hardware exceeds the value of even 20x cloud subscriptions.

    1. zozbot234 · · focus · HN ↗
      Astra-Ultra? Even the largest open model to date (Kimi K3) is nowhere close to Astra level, and it will be quite slow even on the highest-spec M5 Ultra, with achievable speeds of about 0.5 tok/s at most due to having to stream weights from SSD (~13 GB/s on the highest storage capacity M5 Max machines so far). This is OK for doing simple Q&A in the background but it's far from a genuine coding experience. You'd have to test batching of multiple thinking streams in order to try and raise overall tok/s via layer-wise reuse of the streamed weights (and this is where the "Ultra" part sort of becomes relevant; Kimi series models have good support for agent swarms) but this would decrease single-session performance even further. It would only be usable for background jobs, though the hardware would then have a chance of paying for itself if it was fully used on a 24/7 basis.
      1. srcreigh · · focus · HN ↗
        > You'd have to test batching of multiple thinking streams in order to try and raise overall tok/s via layer-wise reuse of the streamed weights

        isn’t this very straightforward to do..? I thought batching for Qwen models is already proven out.

        > but this would decrease single-session performance even further

        Well let’s take Qwen 3.8 27B. Throughput for M3 at 8 agents is 4x compared to single agent. [1]

        It’s really not clear to me that 8 concurrent agents at half speed will be worse task completion latency than 1 agent.

        And that’s M3 studio benchmarks, not even M5 ultra, and without the many software improvements we will see

        If you haven’t tried Qwen 3.8 27B xhigh on a task you might not get the hype. Idk.

        If you’ve tried doing this and don’t like it sure, and be specific about what isn’t effective, but let’s not speculate.

        [1]: <a href="https:&#x2F;&#x2F;omlx.ai&#x2F;benchmarks&#x2F;performance&#x2F;69kzkrv8?utm_source=chatgpt.com" rel="nofollow">https:&#x2F;&#x2F;omlx.ai&#x2F;benchmarks&#x2F;performance&#x2F;69kzkrv8?utm_source=c...

        1. zozbot234 · · focus · HN ↗
          That&#x27;s all well and good but Qwen 27B is a small, dense model; that&#x27;s favorable to both batching and MTP. Batching of large, sparse&#x2F;MoE models like Kimi K3 (requiring slow SSD streaming even on a single maxed out Mac Studio) on local hardware is an entirely different game that&#x27;s mostly theoretical so far: many people would even call it outright pointless. (MTP clearly fares even worse, though - unlike batching, it ends up wasting scarce weights-fetching throughput on wrongly predicted tokens.)
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.