‹ BackHN Continuity

Thread

M5 Ultra Mac Studio Review

269 points · 262 comments · piotrgrabowski

  1. ApolloFortyNine · · focus · HN ↗
    The model being tested is 18k as configured.

    I didn't expect this to make the 5090 to look like a good deal.

    1. nacs · · focus · HN ↗
      5090 has 32GB VRAM.

      It'd be silly to buy the 18k model to run a tiny model like Qwen 27B. You use models like GLM Flash and Qwen Next which won't fit on a single 5090.

      1. asimovDev · · focus · HN ↗
        can run multiple subagents of Qwen 27B though, right? Unless I am fundamentally misunderstanding how VRAM constraints work
        1. Eisenstein · · focus · HN ↗
          You might be. Running another agent doesn't load a set of new weights. It creates a new KV cache for the agent and adds the prompts to the queue. Its just another inference turn.
          1. asimovDev · · focus · HN ↗
            thanks, I naively assumed when, for example, Claude Code starts subagents it loads a new instance with empty context
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.