This is the golden age of model training. Some days ago, I decided I wanted a local CPU only model that can perform exceptionally well for English to Bash translation (to avoid the googling for command syntax). I got a bunch of subagents to generate large amount of training data (140k+ samples), got the Qwen 3 0.6B base model, pointed Astra at it, and off to the races. It trained for 2 days (on and off) and I got a surprisingly good model for my task! The total active time I spent was a few hours. And it is still improving, what a time to be alive!
If it’s one of thing that you want just for English to bash shell commands, I will create AST, it is deterministic, exceptionally fast, no tokens so no need to fine tune existing model, please let me know your thoughts.
You can go quite far using a human language to Bash grammar based setup but at some point the input prompts are harder to translate. The OP has existing projects that work with AST quite deeply so I assume they know about that already.
I am building a natural language to CSV/Excel commands for a "wrangler" type desktop app. Same issues. The MVP is being built with parsers of sorts, entirely code generated. Then I want to fine-tune a tiny model at some point.
GodelNumbering · · focus · HN ↗
teeskay · · focus · HN ↗
brainless · · focus · HN ↗
I am building a natural language to CSV/Excel commands for a "wrangler" type desktop app. Same issues. The MVP is being built with parsers of sorts, entirely code generated. Then I want to fine-tune a tiny model at some point.
<a href="https://github.com/brainless/baho" rel="nofollow">https://github.com/brainless/baho