‹ BackHN Continuity

Thread

Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint

589 points · 200 comments · JonSchneider

  1. danbrooks · · focus · HN ↗
    Nice! Does anyone know how this compares to the Unsloth quantizations of this model? <a href="https:&#x2F;&#x2F;unsloth.ai&#x2F;docs&#x2F;models&#x2F;qwen3.8#run-qwen3.8-guide">https:&#x2F;&#x2F;unsloth.ai&#x2F;docs&#x2F;models&#x2F;qwen3.8#run-qwen3.8-guide
    1. 0xbadcafebee · · focus · HN ↗
      Came to ask the same. From my really rough understanding, it seems like Unsloth&#x27;s method allows a slightly higher precision at a higher file size, while PrismML&#x27;s uses a different approach to achieve a smaller size (and presumably less precision).
    2. nulld3v · · focus · HN ↗
      There&#x27;s a table on the HF page that compares it against Unsloth&#x27;s UD-Q4_K_XL and IQ2_XXS: <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;prism-ml&#x2F;Ternary-Bonsai-2-27B-gguf#full-per-benchmark-results" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;prism-ml&#x2F;Ternary-Bonsai-2-27B-gguf#fu...
      1. kadoban · · focus · HN ↗
        Oh, wow, they think it&#x27;s just a smidge below the q4? That&#x27;s crazy good if true.
        1. anana_ · · focus · HN ↗
          The benchmarks they chose are rather cherry picked to not include long context or difficult ones that involve long horizon work or many agent turns, as I suspect this is where the model shows more differences compared to the full fat one
          1. SillyUsername · · focus · HN ↗
            Yep more hops from the lower Q is likely going to skew the vectors further over time.

            I wonder if there&#x27;s a way to mitigate this by running it through an original Q8 draft model, attuned somehow for the PTQ1 quant, but giving it a higher threshold for the acceptance linear with the context length itself?

            The longer the context, the higher the multiplier on the threshold, and more likely the draft result is used. Not ideal but it may extend the usable max context.

            This model might, even without this, be amazing for short lived agents that work via generations &#x2F; have changing tasks.

    3. [deleted] · · focus · HN ↗

      [deleted]

    4. WithinReason · · focus · HN ↗
      Unsloth has been dethroned by ISTA:

      <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;ISTA-DASLab&#x2F;Qwen3.8-27B-GSQ-RCO-GGUF" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;ISTA-DASLab&#x2F;Qwen3.8-27B-GSQ-RCO-GGUF

      The 3-bit quant is lossless based on benchmarks.

      1. raylad · · focus · HN ↗
        I just checked that model (ISTA-DASLab&#x2F;Qwen3.8-27B-GSQ-RCO-GGUF:IQ3_S) and it does much much worse on the &quot;Please recite Jabberwocky&quot; test than the original bf16 does.

        The bf16 only misses &quot;snicker-snack&quot; and this quantization becomes confused after the first stanza.

        1. sorenjan · · focus · HN ↗
          I think it should be expected that a smaller model is worse at reciting memorized data than a big one. I also don&#x27;t think it&#x27;s a good use case of small local models. Can it find and recite Jabberwocky if given access to a web search tool?
          1. 0xbadcafebee · · focus · HN ↗
            Forgetting things isn&#x27;t lossless though is it? Makes the benchmark and the finding quite suspect
            1. sorenjan · · focus · HN ↗
              It depends on what you mean by lossless. If both models can perform the same tasks it can be considered lossless for those tasks. That task might be more related to language understanding rather than memorizing, they have several benchmarks in the article.
        2. WithinReason · · focus · HN ↗
          I tried and ended up with:

          I&#x27;m going to stop here and be direct: I&#x27;m having trouble recalling the exact text, and every attempt above is me guessing. Rather than present a mangled version as the real poem, I&#x27;d recommend you look it up — it&#x27;s very short and in the public domain, so any text of Through the Looking-Glass will have it verbatim. If you&#x27;d like, I can help with the moral of the poem (&quot;&#x27;twas the blessing of the Bird...&quot;), the famous Humpty Dumpty word interpretations (&quot;slithy&quot; = lithe + sinister, &quot;mimsy&quot; = miserable + mys... etc.), or Carroll&#x27;s original annotations for the coined words — that part I can do reliably.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.