‹ BackHN Continuity

Thread

Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint

589 points · 200 comments · JonSchneider

  1. danbrooks · · focus · HN ↗
    Nice! Does anyone know how this compares to the Unsloth quantizations of this model? <a href="https:&#x2F;&#x2F;unsloth.ai&#x2F;docs&#x2F;models&#x2F;qwen3.8#run-qwen3.8-guide">https:&#x2F;&#x2F;unsloth.ai&#x2F;docs&#x2F;models&#x2F;qwen3.8#run-qwen3.8-guide
    1. WithinReason · · focus · HN ↗
      Unsloth has been dethroned by ISTA:

      <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;ISTA-DASLab&#x2F;Qwen3.8-27B-GSQ-RCO-GGUF" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;ISTA-DASLab&#x2F;Qwen3.8-27B-GSQ-RCO-GGUF

      The 3-bit quant is lossless based on benchmarks.

      1. raylad · · focus · HN ↗
        I just checked that model (ISTA-DASLab&#x2F;Qwen3.8-27B-GSQ-RCO-GGUF:IQ3_S) and it does much much worse on the &quot;Please recite Jabberwocky&quot; test than the original bf16 does.

        The bf16 only misses &quot;snicker-snack&quot; and this quantization becomes confused after the first stanza.

        1. sorenjan · · focus · HN ↗
          I think it should be expected that a smaller model is worse at reciting memorized data than a big one. I also don&#x27;t think it&#x27;s a good use case of small local models. Can it find and recite Jabberwocky if given access to a web search tool?
          1. 0xbadcafebee · · focus · HN ↗
            Forgetting things isn&#x27;t lossless though is it? Makes the benchmark and the finding quite suspect
            1. sorenjan · · focus · HN ↗
              It depends on what you mean by lossless. If both models can perform the same tasks it can be considered lossless for those tasks. That task might be more related to language understanding rather than memorizing, they have several benchmarks in the article.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.