This is the golden age of model training. Some days ago, I decided I wanted a local CPU only model that can perform exceptionally well for English to Bash translation (to avoid the googling for command syntax). I got a bunch of subagents to generate large amount of training data (140k+ samples), got the Qwen 3 0.6B base model, pointed Astra at it, and off to the races. It trained for 2 days (on and off) and I got a surprisingly good model for my task! The total active time I spent was a few hours. And it is still improving, what a time to be alive!
A very small, highly specialized model can use negligible resources (CPU, energy) to accomplish the same task.
For everyday work that happens frequently it's better to have a tiny specialized model instead of making billable API calls or turning your laptop into an 80W space heater for 20 seconds to run a general purpose model.
The large models can be used to generate synthetic training data. Tell them to make up 100,000 tasks paired with the resulting output as a 1-time cost. Then use that to train a small model.
Think of it as distillation, but focused on a specific task.
Given that they're just using it to avoid the googling for bash command syntax, I'm not sure they'll save in the end against the 140k training examples they generated.
You can buy a $10 subscription for a month to generate the training data, then cancel your subscription. The trained model is yours to use (and share with others) forever.
Yea specialized models could also be resold in a shareware style, like even if it cost 50$ to produce you'd just have to sell 10 copies of it to people for 5$.
Which 5$ is a pretty easy sell if its useful in any way, It's pretty easy to justify a purchase if its yours forever and doesn't use much CPU so is easy to run I mean people were spending 1000$+ on mac mini setups to run local llms or run remote agents.
It wouldn't cost $10 for their lifetime use if they used cheap models instead. Mistral Small 3 24B is probably substantially better than their Qwen 3 0.6B trained model and would give you about 500,000–2,000,000 English to Bash translations. If they're a heavy user, they might use $0.30 in total, ever.
They also used Astra for the coding, which they can't get on a $10 subscription. And then there's the actual training cost.
GodelNumbering · · focus · HN ↗
amelius · · focus · HN ↗
Or are the subagents generating your training data using a closed/paid model?
Aurornis · · focus · HN ↗
For everyday work that happens frequently it's better to have a tiny specialized model instead of making billable API calls or turning your laptop into an 80W space heater for 20 seconds to run a general purpose model.
The large models can be used to generate synthetic training data. Tell them to make up 100,000 tasks paired with the resulting output as a 1-time cost. Then use that to train a small model.
Think of it as distillation, but focused on a specific task.
nearbuy · · focus · HN ↗
selcuka · · focus · HN ↗
carsoon · · focus · HN ↗
Which 5$ is a pretty easy sell if its useful in any way, It's pretty easy to justify a purchase if its yours forever and doesn't use much CPU so is easy to run I mean people were spending 1000$+ on mac mini setups to run local llms or run remote agents.
nearbuy · · focus · HN ↗
They also used Astra for the coding, which they can't get on a $10 subscription. And then there's the actual training cost.