Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash
Unofficial Hacker News client; not affiliated with Y Combinator.
viccis · · focus · HN ↗
I have a test suite that tries like ~36 different scenarios, including things like starting multiple timers, saying "actually cancel that timer" and whether it knows to do that one you just created. Basic decision making on top of tool calling. I found so far that, for example, Qwen3.8 on my local machine does pretty poorly even relative to Gemma4 E4B (~9.6gb) and that the best price/performance outcome I've found so far with openrouter is actually GPT Luna, but obviously I'd love to get something that works as well running locally for privacy reasons.
Would love to try this out, I'll just need to tweak my benchmarker to use however this serves it.
HenryNdubuaku · · focus · HN ↗
For your question on tool calling, I think you will find that the model is pretty good at simpler tool calls and parallel ones, but can struggle with implied references and multistep reasoning. These are definitely things that can improve with task-specific finetuning but for some things you just have to have a model that is properly sized. That said, we are always trying to improve the model so that it can handle an ever larger set of queries
viccis · · focus · HN ↗
What it gets wrong:
- It copies numbers instead of converting them. "25 minute timer" becomes duration_seconds: 25, and "twelve minutes" becomes 120. The first one comes with 100% confidence.
- It picks the wrong action. "take the paper towels off the list" became an add. "remind me in 20 minutes" became a timer. "add five minutes to the pasta timer" became a new timer plus a cancel.
- It never declined anything with our full tool set. Background chatter became note_save "blue one" at 0.99 confidence. "play some jazz" became a screen card, and "wake me up at 6 30" became a 630-second timer.
- It can't use household context. Notes, timer names and reminder IDs have no place in its input. Passing them anyway made results worse (5 of 27 single-turn requests right, versus 8 of 27 without), so the backend now leaves them out.
- Follow-ups mostly broke. "take off the last one" removed the whole list.
HenryNdubuaku · · focus · HN ↗