‹ BackHN Continuity

Thread

Ember-1

589 points · 249 comments · gmays

  1. GodelNumbering · · focus · HN ↗
    This is the golden age of model training. Some days ago, I decided I wanted a local CPU only model that can perform exceptionally well for English to Bash translation (to avoid the googling for command syntax). I got a bunch of subagents to generate large amount of training data (140k+ samples), got the Qwen 3 0.6B base model, pointed Astra at it, and off to the races. It trained for 2 days (on and off) and I got a surprisingly good model for my task! The total active time I spent was a few hours. And it is still improving, what a time to be alive!
    1. brainless · · focus · HN ↗
      I have been trying a mix of fine-tuning and I am amazed that most people do not see this coming.

      A tiny, smaller than 1b parameter model, fine-tuned, can kick ass for constrained work. I do not have a lot of budget, I fine-tune only on a 16GB M4 Mac Mini. But that also tells me the potential is wild. Progress has been slow since I moonlight on this.

      I have been trying to build a set of models + agents for full-stack development, where each model does only a small piece, like take user prompt and break into backend/frontend tasks. Then a Rust+Diesel model, a Rust+Auxum model, a Solid+Router model and so on. I know this is wild but this is just theory - can 5 or 6 Qwen 3.5 0.8b models do full-stack web development? My hunch says they can, better than what most people expect. Heck, with a good harness, it might beat all the cheaper models for the specific task, like Haiku or Luna.

      1. jordz · · focus · HN ↗
        I think a sub-set of people see this coming, I also think it isn’t just fine tuning open weight LLM models. A few people I know who are thinking along the same lines with architectures like BERT etc.

        That being said it’s much easier at the moment to continue to use the frontier providers for most general tasks, that is the argument I’ve heard.

        For creating these types of fine tuned local models, on constrained hardware for inference, I do think this is the way to go for specific tasks too!

        1. awwaiid · · focus · HN ↗
          Maybe we're getting into a world where our general models can make their own separate fine-tuned models as utilities just like they would a bash/python script. If making a fine-tune is cheap and fast then it can also be throw-away and constantly improved.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.