‹ BackHN Continuity

Thread

Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash

236 points · 93 comments · HenryNdubuaku

  1. Tsarp · · focus · HN ↗
    Apart from fictional use cases, what is the real use case here? The pricing on some open models are absurdly low for generic tasks. For the privacy conscious it makes sense to run something like a 8-27B on local network and get the work done.

    Are there perhaps some industrial or agri use cases?

    1. HenryNdubuaku · · focus · HN ↗
      Fair, we gotta do a better job at explaining this properly!

      So an 8-27B on a LAN box wins for generic tasks on hardware that can hold it. Needle is for hardware that can't, like plain ARMv7, MIPS32 (the Ingenic chips in cheap IP cameras), RISC-V and watches. Also, we found cost to not really be the lever for on-device models, but availability and latency.

      1. yorwba · · focus · HN ↗
        A list of hardware platforms doesn't make a use case. Do you have an active deployment of Needle that is noticeably useful, and if so, what do you have it do?
        1. stymaar · · focus · HN ↗
          At my current company we're evaluating small models embedded directly in the web app to provide a natural language interface to the app without spending money on inference (and ideally avoid a ChatChipotle situation where people end up having free token going through our interface).

          And needle is one of the most promising model due to its original architecture (but we still need to finish building the actual eval dataset before making out final call).

          1. HenryNdubuaku · · focus · HN ↗
            Thanks for considering needle. Keep in mind that you can also fine-tune the model to fit your use case more. I think this illustrates the intended deployment pretty well, where both computational resources and compute credits can both be issues for deployment.
            1. stymaar · · focus · HN ↗
              We definitely intend to try fine-tuning, don't worry we're not going to dismiss needle just because the base model's performance is too low ;).
          2. int_19h · · focus · HN ↗
            You can embed much bigger models into web apps with wasm and WebGL or WebGPU. I have a web app running a 0.6B embedding model client-side.
            1. stymaar · · focus · HN ↗
              WebGPU is a non-starter for production usage since the support is too limited (No Firefox support, no Linux Support, no Apple x86 support, no MacOS <26).

              But yes we are also considering bigger models, though we'll pick the smallest model of sufficient quality because not having to download a 600MB bag of weight is a feature in itself.

              1. int_19h · · focus · HN ↗
                That very much depends on your target audience (but also Firefox has partial support these days). However WebGL is widely available and can be used as a fallback - it might be less efficient but for models this small it can still be acceptable.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.