‹ BackHN Continuity

Thread

Ember-1

589 points · 249 comments · gmays

  1. GodelNumbering · · focus · HN ↗
    This is the golden age of model training. Some days ago, I decided I wanted a local CPU only model that can perform exceptionally well for English to Bash translation (to avoid the googling for command syntax). I got a bunch of subagents to generate large amount of training data (140k+ samples), got the Qwen 3 0.6B base model, pointed Astra at it, and off to the races. It trained for 2 days (on and off) and I got a surprisingly good model for my task! The total active time I spent was a few hours. And it is still improving, what a time to be alive!
    1. amelius · · focus · HN ↗
      I don't understand. If you have a model that can do bash examples already (your subagents), then why would you need to train a model?

      Or are the subagents generating your training data using a closed/paid model?

      1. Aurornis · · focus · HN ↗
        A very small, highly specialized model can use negligible resources (CPU, energy) to accomplish the same task.

        For everyday work that happens frequently it's better to have a tiny specialized model instead of making billable API calls or turning your laptop into an 80W space heater for 20 seconds to run a general purpose model.

        The large models can be used to generate synthetic training data. Tell them to make up 100,000 tasks paired with the resulting output as a 1-time cost. Then use that to train a small model.

        Think of it as distillation, but focused on a specific task.

        1. nearbuy · · focus · HN ↗
          Given that they're just using it to avoid the googling for bash command syntax, I'm not sure they'll save in the end against the 140k training examples they generated.
          1. verdverm · · focus · HN ↗
            you can probably generate quite a few example pairs in a single shot, you also likely don't need the best models for this either
            1. liquicity · · focus · HN ↗
              Same token burn / cost though right?

              I think OP's point remains, if you generate 140k pairs, your local model would need to run that many to offset having just used the generator (SOTA or not) model to begin with.

              I wonder if another approach if latency is a concern is just to do a two shot pass with Jev (perhaps given small context you'd want one to match command, then one to match args of given command) would be an extremely fast, and cheap way to do it - rather than training your own.

              1. viraptor · · focus · HN ↗
                > Same token burn / cost though right?

                Not quite. One, because you save on the initial query being sent multiple times. Two, because the reasoning will be very similar at the beginning; "user asked me to", "let's check what's in this project already", etc. you'll get similar actual output cost, but input and reasoning will be shorter.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.