‹ BackHN Continuity

Thread

M5 Ultra Mac Studio Review

269 points · 262 comments · piotrgrabowski

  1. srcreigh · · focus · HN ↗
    This is great as a first look, but the author is not a developer, so we don't yet know whether a dev can be as productive with local models on M5 Mac Studio compared to a 20x subscription plan.

    I'm also curious about any new low hanging optimization opportunities in the kernels for this new hardware.

    It's already clear to me that M5 Mac Studio is more cost-effective than anything you can run on open router, assuming decent utilization.

    The M5 Mac Studio will be the most cost effective way to run uncensored cyber capable open agents.

    An exciting tipping point will be if programmers can get an Astra-Ultra like experience all week with this hardware. That would be a real sense where this hardware exceeds the value of even 20x cloud subscriptions.

    1. slowin · · focus · HN ↗
      > This is great as a first look, but the author is not a developer, so we don't yet know whether a dev can be as productive with local models on M5 Mac Studio compared to a 20x subscription plan.

      Local models are definitely not as productive as SOTA, sadly it's not close yet. I do think someday they will be "good enough" to use, but they aren't today. Even the SOTA models barely code well, with Opus 4.5 being the first, good coding model.

      That being said, I think it's absolutely imperative that we keep pushing local model performance. We need to continue to advance technology there and ensure that the model labs don't do regulatory capture in the name of "safety" (or anything else).

      1. _hugerobots_ · · focus · HN ↗
        Local models can be widely used as productive assets. Yes the infrastructure of SOTA API models is engineered specifically for you to be that utility, but the blanket statement that local isn't up to par is intensely short sighted. Billions of tokens per month on local pays for the hardware when compared to sota costs per month.
        1. slowin · · focus · HN ↗
          I believe they can currently be used productively for non-coding tasks (classification, light summary)... but they definitely are not even close to SOTA when it comes to software development.
          1. _hugerobots_ · · focus · HN ↗
            Defining productivity is a use-case scenario, and a wildly generalized assumption for most people in this argument. Local infrastructure doesn't need to be sota for absolutely every single need for a dev lab, but it absolutely can be delivered with non-api frontier class models.
            1. slowin · · focus · HN ↗
              Just to be clear, I'm specifically talking about coding. I think local models can help with productivity today, just not coding.

              I'm also a huge fan of local models and think it's absolutely imperative that they continue to advance so we can move off of the Anthropic/OpenAI hosted models. It's important to accurately asses where we are in that journey though.

              1. srcreigh · · focus · HN ↗
                I think the issue is generalization, if you were more specific about which local models aren’t good enough for which tasks compared to which frontier models in your experience, it’d be a lot more informative
                1. slowin · · focus · HN ↗
                  I can't just go into any codebase and ask a local model to "Implement this feature: xxx" and get acceptable output. I hope to someday soon though!
              2. _hugerobots_ · · focus · HN ↗
                Like the other commenter, I'm confused about the 'just not coding' conclusion. I'm using Qwen 27B on a 5090 at > 100tk/s with 150k context (which isn't enough admittedly), and DeepSeek v4 Flash with 1million context on a gb10/spark. Both of which are performing surface level, and deep needle precision infrastructure architecture. They code 24-7, stupendously.
                1. fhub · · focus · HN ↗
                  It would be interesting to hear more about how you’re actually using them. Do you have sophisticated feedback loops around the models so they can verify their work and converge on good solutions? And how do you decide what to give the 5090 vs the Spark vs a frontier model?

                  Correctness matters much more than speed to me, but if I can get both, that’s obviously very interesting.

              3. brandon272 · · focus · HN ↗
                Local models are undeniably capable of "helping with coding" today.
                1. slowin · · focus · HN ↗
                  I so want this to be true, but for the kind of coding I do (not Flask apps), it's definitely not the case. Like I said, SOTA models just barely, barely work for me. My projects are usually 100k-1M lines of Rust or Go.
                  1. brandon272 · · focus · HN ↗
                    Out of curiosity, what do you find the SOTA models are simply incapable of when it comes to your Rust and Go projects?
                    1. slowin · · focus · HN ↗
                      The SOTA models now work really well in my codebases, but that's only been since Opus 4.5/4.6-ish. Prior to that, and with current local models, they simply couldn't work holistically and would just thrash around. Now I feel as if SOTA are approaching my coding levels if not surpassing it. I still need to guide on architecture, but I can see that going away within the next year or so as well.
                      1. brandon272 · · focus · HN ↗
                        Thanks, that makes sense. When you said they “barely, barely worked” for you I assumed that meant something different.
                        1. slowin · · focus · HN ↗
                          Oh yeah, that makes sense, sorry! I meant they just started working well and did not until relatively recently.
              4. poincareball · · focus · HN ↗

                [dead]

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.