‹ BackHN Continuity

Thread

Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint

589 points · 200 comments · JonSchneider

  1. nilsherzig · · focus · HN ↗
    Fyi, if you're trying to run this under AMD/HIP:

    PTQ1_0 has no optimized MMQ-Path in their llama-cpp fork, try running PTQ2_0 (needs a bit more vram, but is about 2x faster on my 6700 XT)

    <a href="https:&#x2F;&#x2F;gist.github.com&#x2F;nilsherzig&#x2F;b8266d001c5c01bdb3d81d20915572fc" rel="nofollow">https:&#x2F;&#x2F;gist.github.com&#x2F;nilsherzig&#x2F;b8266d001c5c01bdb3d81d209...

    1. jakswa · · focus · HN ↗
      what kinda speeds do you see on 6700 XT? i&#x27;m always conflicted on investing time chasing speed-vs-quality tradeoffs. I&#x27;ve got a 7900 XT (about double the IO throughput). I&#x27;ll probably end up giving it a go when I find time.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.