‹ BackHN Continuity

Thread

Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint

589 points · 200 comments · JonSchneider

  1. huseyinkeles · · focus · HN ↗
    Testing on a MBP m4 pro 24gb

    ~100t/s prefill, ~15t/s, dropping to ~10t/s later with 64k context.

    The issue is I have yet to find a useful agentic local llm that I can run on this machine.

    Just given a relatively simple task on a swift app, took 25 minutes, brainstorming like crazy but can not decide on what to do. Eventually I killed it. GPT 5.6 sol-medium took 3 minutes to complete the same task for reference.

    1. sean_pedersen · · focus · HN ↗
      Try a MoE model like Qwen3.6 35B-A3B for better tok/s
      1. huseyinkeles · · focus · HN ↗
        I tried this one, but I found Ornith1.5 to be a better MoE model for me, also very fast. But I still couldn't make it implement a real task on a real repo :( it only worked with an extremely clear directions and very small tasks.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.