‹ BackHN Continuity

Thread

Shapelearn Qwen 3.8 27B (13.1 GB VRAM)

104 points · 39 comments · syntaxing

  1. npodbielski · · focus · HN ↗
    Well I tested it on 7900XTX with the same prompts and their draft model gave me about 30t/s. Their own snippet of code with regular MTP model gave me 60t/s.

    Also model with their draft answered incorrectly. With MTP it answered correctly.

    Question was: "Does MikroTik CRS312-4C+8XG-RM have combo ports?". The answer is Yes.

    1. nvme0n1p1 · · focus · HN ↗
      That's not how you're supposed to use LLMs. You shouldn't expect a tiny little local model to know random facts about every obscure consumer product on earth. That's the job of tool calling. At best a model of this size is just giving you a random guess.

      You're basically saying "I tried rolling these dice one time, the green dice rolled a 6 and the blue dice rolled a 1, so green dice are better"

      1. serf · · focus · HN ↗
        agreed. a niche knowledge callout is about the worst benchmark one can give a smaller model.

        smaller models are attempting to distill the useful methodologies, not the license plate number of an obscure extras car on Magnum PI.

        that said I wonder if there is a small 'trivia' model out there. Seems like the kinda thing Google would tackle.

      2. npodbielski · · focus · HN ↗
        Which was not he point because I was testing their solution for MPT and it was just funny addition. But of course in internet you always will find some 'well akchually' person straight from the meme.
    2. npodbielski · · focus · HN ↗
      When I changed the number of draft tokens to 3 in both, it helped and they Draft is actually performing a bit better:

      - draft: 67.17

      - MTP: 64.18

      Why they used those examples? Seems strange.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.