‹ BackHN Continuity

Thread

Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint

589 points · 200 comments · JonSchneider

  1. jakswa · · focus · HN ↗
    Bonsai 2 27B · Radeon RX 7900 XTX

      - 89 tokens/sec generation with speculative decoding
      - 81 tokens/sec at 20k context
      - 474 tokens/sec ingestion at 20k — about 42 seconds
      - 10.1 GiB peak VRAM with a 24k context window
    
    ROCm 7.2.3 · PQ2_0 · Qwen Q4 MTP, draft length 2

    ---- versus ----

    Qwen3.8-27B IQ3_S · Radeon RX 7900 XTX

      - 79 tokens/sec generation on a short coding prompt
      - 61 tokens/sec at 60k context
      - 53 tokens/sec at 95k context
      - 558 tokens/sec ingestion at 60k — about 108 seconds
      - 19.9 GiB peak VRAM during coding tests with a 100k context window
    
    Vulkan · GSQ-RCO IQ3_S · MTP, draft length 2 · vision projector loaded
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.