‹ BackHN Continuity

Thread

Shapelearn Qwen 3.8 27B (13.1 GB VRAM)

104 points · 39 comments · syntaxing

  1. kristianp · · focus · HN ↗
    What's GPU-5?
    1. LtdJorge · · focus · HN ↗
      The fattest quantization. They show all of them here: <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;byteshape&#x2F;Qwen3.8-27B-GGUF" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;byteshape&#x2F;Qwen3.8-27B-GGUF

      Since they’re not using a stable number of bits per token, they use their own naming convention.

      1. noir_lord · · focus · HN ↗
        &gt; they use their own naming convention.

        Seems like a lot of them do, I only compare them within the same repo because there doesn&#x27;t seem to be a very standard way of saying all the possible combinations&#x2F;rearrangements.

        1. iker00 · · focus · HN ↗
          bpw is the way to compare. huggingface has standard tags that must be used so it forces anyone releasing models to choose a tag that doesn&#x27;t necessarily equal the actual bpw.
          1. tancop · · focus · HN ↗
            It&#x27;s a bad way to compare. Average bpw ignores the fact that some layers can tolerate more aggressive quant than others.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.