‹ BackHN Continuity

Thread

M5 Ultra Mac Studio Review

269 points · 262 comments · piotrgrabowski

  1. ApolloFortyNine · · focus · HN ↗
    The model being tested is 18k as configured.

    I didn't expect this to make the 5090 to look like a good deal.

    1. nacs · · focus · HN ↗
      5090 has 32GB VRAM.

      It'd be silly to buy the 18k model to run a tiny model like Qwen 27B. You use models like GLM Flash and Qwen Next which won't fit on a single 5090.

      1. orsorna · · focus · HN ↗
        Is it that silly? You could run multiple 27B models in parallel.
        1. peri-cl · · focus · HN ↗
          You actually don't need more RAM to batch multiple inference tasks of the same model.

          (Each task needs its own context, but the (e.g.) 27B of constant parameters isn't duplicated).

          1. orsorna · · focus · HN ↗
            You definitely need more RAM if you are not satisfied with small context windows, especially if the weights take a large % of the total memory to boot.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.