‹ BackHN Continuity

Thread

Ember-1

589 points · 249 comments · gmays

  1. GodelNumbering · · focus · HN ↗
    This is the golden age of model training. Some days ago, I decided I wanted a local CPU only model that can perform exceptionally well for English to Bash translation (to avoid the googling for command syntax). I got a bunch of subagents to generate large amount of training data (140k+ samples), got the Qwen 3 0.6B base model, pointed Astra at it, and off to the races. It trained for 2 days (on and off) and I got a surprisingly good model for my task! The total active time I spent was a few hours. And it is still improving, what a time to be alive!
    1. amelius · · focus · HN ↗
      I don't understand. If you have a model that can do bash examples already (your subagents), then why would you need to train a model?

      Or are the subagents generating your training data using a closed/paid model?

      1. Aurornis · · focus · HN ↗
        A very small, highly specialized model can use negligible resources (CPU, energy) to accomplish the same task.

        For everyday work that happens frequently it's better to have a tiny specialized model instead of making billable API calls or turning your laptop into an 80W space heater for 20 seconds to run a general purpose model.

        The large models can be used to generate synthetic training data. Tell them to make up 100,000 tasks paired with the resulting output as a 1-time cost. Then use that to train a small model.

        Think of it as distillation, but focused on a specific task.

        1. nearbuy · · focus · HN ↗
          Given that they're just using it to avoid the googling for bash command syntax, I'm not sure they'll save in the end against the 140k training examples they generated.
          1. verdverm · · focus · HN ↗
            you can probably generate quite a few example pairs in a single shot, you also likely don't need the best models for this either
            1. liquicity · · focus · HN ↗
              Same token burn / cost though right?

              I think OP's point remains, if you generate 140k pairs, your local model would need to run that many to offset having just used the generator (SOTA or not) model to begin with.

              I wonder if another approach if latency is a concern is just to do a two shot pass with Jev (perhaps given small context you'd want one to match command, then one to match args of given command) would be an extremely fast, and cheap way to do it - rather than training your own.

              1. viraptor · · focus · HN ↗
                > Same token burn / cost though right?

                Not quite. One, because you save on the initial query being sent multiple times. Two, because the reasoning will be very similar at the beginning; "user asked me to", "let's check what's in this project already", etc. you'll get similar actual output cost, but input and reasoning will be shorter.

          2. kubb · · focus · HN ↗
            Good observation! It would have to be offset with O(140k) queries to the model, which is, well, unlikely.
            1. computably · · focus · HN ↗
              If it's about the latency / flow disruption, spending a few hours once could easily be worth it if the result is actually good enough to skip googling/retries.
            2. Lalabadie · · focus · HN ↗
              Just like with OSS in general, being able to distribute it is what makes the effort worthwhile.

              This particular example is maybe a niche, but 1400 people can use a few hundred queries in a reasonable amount of time.

              1. 8n4vidtmkvmk · · focus · HN ↗
                This example is not that niche. Lots of people use human to bash. I'd probably use a small pre trained model if it was easy to use. I use a little script right now that calls a cheap model. Actually.. $5 would probably will last me over a year so the only real benefit would be if I didn't have an internet connection.
          3. selcuka · · focus · HN ↗
            You can buy a $10 subscription for a month to generate the training data, then cancel your subscription. The trained model is yours to use (and share with others) forever.
            1. carsoon · · focus · HN ↗
              Yea specialized models could also be resold in a shareware style, like even if it cost 50$ to produce you'd just have to sell 10 copies of it to people for 5$.

              Which 5$ is a pretty easy sell if its useful in any way, It's pretty easy to justify a purchase if its yours forever and doesn't use much CPU so is easy to run I mean people were spending 1000$+ on mac mini setups to run local llms or run remote agents.

            2. nearbuy · · focus · HN ↗
              It wouldn't cost $10 for their lifetime use if they used cheap models instead. Mistral Small 3 24B is probably substantially better than their Qwen 3 0.6B trained model and would give you about 500,000–2,000,000 English to Bash translations. If they're a heavy user, they might use $0.30 in total, ever.

              They also used Astra for the coding, which they can't get on a $10 subscription. And then there's the actual training cost.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.