‹ BackHN Continuity

Thread

Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint

589 points · 200 comments · JonSchneider

  1. kamranjon · · focus · HN ↗
    Love this for the folks with 16gb graphics cards - 3.8 27b has been incredible but not quite runnable on anything less than 32gb - will try loading this up on my 16gb intel b50 and see how it goes - not sure these quants can be accelerated by the XPU cores yet but maybe in time!
    1. kadoban · · focus · HN ↗
      You can run the ~4 bit quant(s) on 24gb, if you're not _too_ picky on context size.

      This will hopefully be better, though it'd be a _very_ surprising increase in performace at the size they say. Would love to see more about how it benchmarks.

      1. lta · · focus · HN ↗
        I'm doing the same with a context of about 128-150k Surprisingly, I get subjectively better results with Unsloth's 3 bit quants (UD-Q3-XL something), than their 4 bit quants (S or M)
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.