‹ BackHN Continuity

Thread

Shapelearn Qwen 3.8 27B (13.1 GB VRAM)

104 points · 39 comments · syntaxing

  1. _ache_ · · focus · HN ↗
    From my own test. It's not faster than the unsloth model.

    Disclarer: I'm unsing Vulkan on an AMD GC.

    1. Systemerror7A69 · · focus · HN ↗
      AMD 7900 XTX with Vulkan here as well, wasn't faster on my test either. Might be much different on Nvidia though.

      I assume the limit for me is memory bandwith, as the 7900 XTX has the same bandwith as the 3090 from what I can gather and I already reached ~60 t/s with Unsloth. Those would fit with the numbers Byteshape has for their cards.

      4090 and 5090 have much higher bandwith apparently, so on those cards you can probably get much more out of the kinds of performance improvements they are doing.

      1. noir_lord · · focus · HN ↗
        4090 isn't that much higher than the XTX (I also have the XTX), it's 1008GB/s (4090) vs 960GB/s for the XTX's.

        The 5090 destroys both at 1792GB/s.

        It's not really one thing with the nvidia cards best I can tell it's that they compounded incremental gains from software drivers, card kernels and optimization from been the primary choice (plus first mover advantage).

        I didn't buy the XTX for AI purely gaming but it's a capable enough local card for running Qwen et al.

        1. Figs · · focus · HN ↗
          4090 vs 5090 performance difference is largely GDDR6 vs GDDR7, I think
          1. wtallis · · focus · HN ↗
            It's equal parts memory clock and bus width: RTX 4090 is 384-bit wide at 21 Gb/s and RTX 5090 is 512-bit wide at 28 Gb/s.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.