‹ BackHN Continuity

Thread

Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash

236 points · 93 comments · HenryNdubuaku

  1. viccis · · focus · HN ↗
    I'll try to get something set up to try this out. I've been working on an ESP32 based Echo replacement that sends audio back to a backend server I run, and one question I had was whether models small enough to run on a Mac Mini or even smaller hardware are good enough to handle basic tool calling functionality with a bit of reasoning where needed.

    I have a test suite that tries like ~36 different scenarios, including things like starting multiple timers, saying "actually cancel that timer" and whether it knows to do that one you just created. Basic decision making on top of tool calling. I found so far that, for example, Qwen3.8 on my local machine does pretty poorly even relative to Gemma4 E4B (~9.6gb) and that the best price/performance outcome I've found so far with openrouter is actually GPT Luna, but obviously I'd love to get something that works as well running locally for privacy reasons.

    Would love to try this out, I'll just need to tweak my benchmarker to use however this serves it.

    1. akadeb · · focus · HN ↗
      hey viccis the tool calling on ESP32 with a Mac M4 is exactly what I built here. check it out and lmk if it helps <a href="https:&#x2F;&#x2F;github.com&#x2F;akdeb&#x2F;open-toys" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;akdeb&#x2F;open-toys
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.