‹ BackHN Continuity

Thread

Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint

589 points · 200 comments · JonSchneider

  1. flutetornado · · focus · HN ↗
    GPT Astra did some benchmarking on the DGX Spark. Speed: 34.38 tokens/sec for generation.

    Seems like we don't have a drafter model yet so it could not test with speculative decoding on. ngram speculative decoding did not help too much either - not enough accepted tokens.

    Smaller size I suppose does not mean better performance in this case - we maybe limited by Spark's low memory bandwidth.

    1. cmrdporcupine · · focus · HN ↗
      What are you getting for prefill?
      1. flutetornado · · focus · HN ↗
        450 with PTQ_01 and 900 with the other PQ2_0.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.