This is the golden age of model training. Some days ago, I decided I wanted a local CPU only model that can perform exceptionally well for English to Bash translation (to avoid the googling for command syntax). I got a bunch of subagents to generate large amount of training data (140k+ samples), got the Qwen 3 0.6B base model, pointed Astra at it, and off to the races. It trained for 2 days (on and off) and I got a surprisingly good model for my task! The total active time I spent was a few hours. And it is still improving, what a time to be alive!
A very small, highly specialized model can use negligible resources (CPU, energy) to accomplish the same task.
For everyday work that happens frequently it's better to have a tiny specialized model instead of making billable API calls or turning your laptop into an 80W space heater for 20 seconds to run a general purpose model.
The large models can be used to generate synthetic training data. Tell them to make up 100,000 tasks paired with the resulting output as a 1-time cost. Then use that to train a small model.
Think of it as distillation, but focused on a specific task.
Given that they're just using it to avoid the googling for bash command syntax, I'm not sure they'll save in the end against the 140k training examples they generated.
This example is not that niche. Lots of people use human to bash.
I'd probably use a small pre trained model if it was easy to use. I use a little script right now that calls a cheap model.
Actually.. $5 would probably will last me over a year so the only real benefit would be if I didn't have an internet connection.
GodelNumbering · · focus · HN ↗
amelius · · focus · HN ↗
Or are the subagents generating your training data using a closed/paid model?
Aurornis · · focus · HN ↗
For everyday work that happens frequently it's better to have a tiny specialized model instead of making billable API calls or turning your laptop into an 80W space heater for 20 seconds to run a general purpose model.
The large models can be used to generate synthetic training data. Tell them to make up 100,000 tasks paired with the resulting output as a 1-time cost. Then use that to train a small model.
Think of it as distillation, but focused on a specific task.
nearbuy · · focus · HN ↗
kubb · · focus · HN ↗
Lalabadie · · focus · HN ↗
This particular example is maybe a niche, but 1400 people can use a few hundred queries in a reasonable amount of time.
8n4vidtmkvmk · · focus · HN ↗