‹ BackHN Continuity

Thread

Ember-1

589 points · 249 comments · gmays

  1. GodelNumbering · · focus · HN ↗
    This is the golden age of model training. Some days ago, I decided I wanted a local CPU only model that can perform exceptionally well for English to Bash translation (to avoid the googling for command syntax). I got a bunch of subagents to generate large amount of training data (140k+ samples), got the Qwen 3 0.6B base model, pointed Astra at it, and off to the races. It trained for 2 days (on and off) and I got a surprisingly good model for my task! The total active time I spent was a few hours. And it is still improving, what a time to be alive!
    1. shriphani · · focus · HN ↗
      what hardware are you using to train?
      1. GodelNumbering · · focus · HN ↗
        I didn't have a local GPU, so I asked it to go out and find hardware. It found a google TPU v6e which seemed reasonably priced. I gave it my google api key. I told it to use TPU only when training and bring it down afterwards. That's about it.
        1. shriphani · · focus · HN ↗
          Neat!
        2. otterley · · focus · HN ↗
          What kind of observability did you have over this process? I’m interested in how my peers are operating these efforts.
          1. GodelNumbering · · focus · HN ↗
            On the cloud side, nothing valuable existed, so the training couldn't ruin anything it didn't create. On the laptop side, I usually ask the agents to create named scripts for everything it needs to access, then those local script directory is green-lit with approve all. For cost, I kept giving it new budget in the 20-30 dollar increments.

            I had to intervene a few times. For instance, as smart as the models are said to be (Astra), it would copy the full training run, train on the server, pull every checkpoint to the local machine, then run tests, update. So, the bandwidth bill was as high as training bill for the first 6 hours. It could have simply tested each checkpoint on the server, saved time and money, didn't occur to it until I said.

            1. otterley · · focus · HN ↗
              Perhaps I wasn’t clear. What kind of instrumentation and alerting, if any, did you employ to keep an eye on it?
              1. GodelNumbering · · focus · HN ↗
                instrumentation: scripts to watch the runs, measure, report. alerting: none, the model was access limited and constrained by other means.
        3. [deleted] · · focus · HN ↗

          [deleted]

        4. varispeed · · focus · HN ↗
          > I told it to use TPU only when training and bring it down afterwards.

          I wouldn't put my house on it. Brave.

        5. libria · · focus · HN ↗
          > I gave it my google api key

          This is the part where the narrator looks at the camera and says "Don't try this at home, kids!"

          1. bitpush · · focus · HN ↗
            Why? Isnt the API key scoped to a project and specifically made for this?

            Are you confusing this with an OAuth token or something?

            1. raizer88 · · focus · HN ↗
              Until astra goes bonkers and use the tpu for days
              1. verdverm · · focus · HN ↗
                this is what billing caps are for

                <a href="https:&#x2F;&#x2F;docs.cloud.google.com&#x2F;billing&#x2F;docs&#x2F;how-to&#x2F;budgets-spend-caps" rel="nofollow">https:&#x2F;&#x2F;docs.cloud.google.com&#x2F;billing&#x2F;docs&#x2F;how-to&#x2F;budgets-sp...

                1. degamad · · focus · HN ↗
                  Billing caps that Google notoriously doesn&#x27;t enforce?
              2. steve_adams_86 · · focus · HN ↗
                This is kind of what it was supposed to do, in this case
          2. edot · · focus · HN ↗
            There’s a safer way to do this with nearly no added friction. Give it a read only API key. Then just ask it to write the API calls into a bash script and then read it and run it yourself. The agent can still inspect the live resources and diagnose and give you more commands to run. I do agree I wouldn’t give it create &#x2F; write access.
          3. MisterMunchkin · · focus · HN ↗
            You’re absolutely right, I shouldn’t have rented a 200 GPU cluster for $35,000&#x2F;hour. That’s on me.

            [Search: Can I refund Google cloud?]

            It looks like we’re not able to ask for a refund since we did actually use all of that compute intentionally.

            Would you like me to write you a pleading email to send to the support team?

          4. hgoel · · focus · HN ↗
            I&#x27;ve done this sort of thing before but with Vast. Pre-deposited some money online, then let the LLM request and manage a training run on an allocation. Worked pretty well without risking bankruptcy.
          5. lopsotronic · · focus · HN ↗
            Oh yeah. Listen to this guy please.
    2. PEe9bB7D · · focus · HN ↗
      i also need more info!
      1. GodelNumbering · · focus · HN ↗
        I am thinking about opensourcing everything, although this is not my main domain or my main startup, so the overhead of huggingface etc seems a bit unnecessary

        Edit: will do as soon as possible

        1. jack_pp · · focus · HN ↗
          just ask the agent to write it up if you don&#x27;t have time to do a write-up yourself
          1. genxy · · focus · HN ↗
            Just use the post-one-off-project-to-huggingface-skill.md
        2. equinumerous · · focus · HN ↗
          +1, would like to see. Even if it&#x27;s not fully &quot;ready for consumption&quot;, it&#x27;s probably enough to reproduce the results.
        3. jjice · · focus · HN ↗
          Please do! Small, specialized models need more love and the time you spent would be a gift!
        4. atombender · · focus · HN ↗
          Would also love to read a write-up about this!
        5. onel · · focus · HN ↗
          Please do share it.
    3. luisfmh · · focus · HN ↗
      Curious about how you generated the training data? Was it just asking an existing model to generate a bunch of examples?

      I ask cause would this be a kind of model distillation?

      I have a small model I&#x27;m looking to train on some data, and I have some real live data but I&#x27;d love to be able to extend it.

      1. GodelNumbering · · focus · HN ↗
        All synthetic data. For this usecase, it was easier because all current generation LLMs, even the small models, are really good at bash commands (and SQL queries too)), so you can reasonably start batches of cheap subagents whose output is reviewed by a more capable model and merge into main training set. After 100k, I had to standing instructions to run the generation loops selectively, meaning only update samples in a given area where we see poor capability.
        1. verdverm · · focus · HN ↗
          Do you have a write-up or git repo for this? Would love to learn more and&#x2F;or dig into the guts

          edit: others have asked any you have replied &quot;soon (tm)&quot;, looking forward for that day

        2. jeremyjh · · focus · HN ↗
          It would be awesome to share your training set on hugging face if it’s easy to de-personalize it. The largest I could find was only 800 rows.
          1. mtud · · focus · HN ↗
            I’ve previously fine tuned (SFT, currently trying to distill) models using this <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;datasets&#x2F;westenfelder&#x2F;NL2SH-ALFA" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;datasets&#x2F;westenfelder&#x2F;NL2SH-ALFA
      2. toasty228 · · focus · HN ↗
        It is a form of distillation, as long as you&#x27;re working a very narrow &quot;trivial&quot; topics it works perfectly.
    4. equinumerous · · focus · HN ↗
      That&#x27;s a really impressive result. There are all kinds of small tasks like this I use an LLM for, but theoretically if you broke all the sub-use cases into local-only models, and had something lightweight that routed to the right model, you could have faster and cheaper workflows. E.g. something trained on the linux man pages for common commands, since it&#x27;s usually quicker to ask an LLM for a specific command with flags than to consult the man pages.
      1. dominotw · · focus · HN ↗
        &gt; That&#x27;s a really impressive result.

        we dont know what the result is and how its impressive.

    5. amelius · · focus · HN ↗
      I don&#x27;t understand. If you have a model that can do bash examples already (your subagents), then why would you need to train a model?

      Or are the subagents generating your training data using a closed&#x2F;paid model?

      1. computerex · · focus · HN ↗
        The models he is using to generate training data are presumably commercial models. He is distilling their bash knowledge into a much smaller model he can run locally fast and cheap.
      2. Aurornis · · focus · HN ↗
        A very small, highly specialized model can use negligible resources (CPU, energy) to accomplish the same task.

        For everyday work that happens frequently it&#x27;s better to have a tiny specialized model instead of making billable API calls or turning your laptop into an 80W space heater for 20 seconds to run a general purpose model.

        The large models can be used to generate synthetic training data. Tell them to make up 100,000 tasks paired with the resulting output as a 1-time cost. Then use that to train a small model.

        Think of it as distillation, but focused on a specific task.

        1. nearbuy · · focus · HN ↗
          Given that they&#x27;re just using it to avoid the googling for bash command syntax, I&#x27;m not sure they&#x27;ll save in the end against the 140k training examples they generated.
          1. verdverm · · focus · HN ↗
            you can probably generate quite a few example pairs in a single shot, you also likely don&#x27;t need the best models for this either
            1. liquicity · · focus · HN ↗
              Same token burn &#x2F; cost though right?

              I think OP&#x27;s point remains, if you generate 140k pairs, your local model would need to run that many to offset having just used the generator (SOTA or not) model to begin with.

              I wonder if another approach if latency is a concern is just to do a two shot pass with Jev (perhaps given small context you&#x27;d want one to match command, then one to match args of given command) would be an extremely fast, and cheap way to do it - rather than training your own.

              1. viraptor · · focus · HN ↗
                &gt; Same token burn &#x2F; cost though right?

                Not quite. One, because you save on the initial query being sent multiple times. Two, because the reasoning will be very similar at the beginning; &quot;user asked me to&quot;, &quot;let&#x27;s check what&#x27;s in this project already&quot;, etc. you&#x27;ll get similar actual output cost, but input and reasoning will be shorter.

          2. kubb · · focus · HN ↗
            Good observation! It would have to be offset with O(140k) queries to the model, which is, well, unlikely.
            1. computably · · focus · HN ↗
              If it&#x27;s about the latency &#x2F; flow disruption, spending a few hours once could easily be worth it if the result is actually good enough to skip googling&#x2F;retries.
            2. Lalabadie · · focus · HN ↗
              Just like with OSS in general, being able to distribute it is what makes the effort worthwhile.

              This particular example is maybe a niche, but 1400 people can use a few hundred queries in a reasonable amount of time.

              1. 8n4vidtmkvmk · · focus · HN ↗
                This example is not that niche. Lots of people use human to bash. I&#x27;d probably use a small pre trained model if it was easy to use. I use a little script right now that calls a cheap model. Actually.. $5 would probably will last me over a year so the only real benefit would be if I didn&#x27;t have an internet connection.
          3. selcuka · · focus · HN ↗
            You can buy a $10 subscription for a month to generate the training data, then cancel your subscription. The trained model is yours to use (and share with others) forever.
            1. carsoon · · focus · HN ↗
              Yea specialized models could also be resold in a shareware style, like even if it cost 50$ to produce you&#x27;d just have to sell 10 copies of it to people for 5$.

              Which 5$ is a pretty easy sell if its useful in any way, It&#x27;s pretty easy to justify a purchase if its yours forever and doesn&#x27;t use much CPU so is easy to run I mean people were spending 1000$+ on mac mini setups to run local llms or run remote agents.

            2. nearbuy · · focus · HN ↗
              It wouldn&#x27;t cost $10 for their lifetime use if they used cheap models instead. Mistral Small 3 24B is probably substantially better than their Qwen 3 0.6B trained model and would give you about 500,000–2,000,000 English to Bash translations. If they&#x27;re a heavy user, they might use $0.30 in total, ever.

              They also used Astra for the coding, which they can&#x27;t get on a $10 subscription. And then there&#x27;s the actual training cost.

        2. jamienk · · focus · HN ↗
          This is so cool - I&#x27;m aware of this in a vague way. Can you write a little tutorial or give some good links. I want this to be the next new things I do :)
          1. newswasboring · · focus · HN ↗
            Better yet package it up in a skill!
        3. tomrod · · focus · HN ↗
          I feel like we need a good index for these kinds of specialized models, especially if you plan to open them up. The downside is a new bash version means potentially new training.
          1. msdz · · focus · HN ↗
            I took it more to mean “do standard Unix command-line stuff” rather than purely emitting Bash, and while e.g. POSIX does receive updates, the basic stuff using the main tools should also work in the future!
        4. askl · · focus · HN ↗
          Or you could just google the syntax to accomplish the same task faster and with fewer resources.
          1. edschofield · · focus · HN ↗
            The same Google that drowns you in low-quality AI responses, ads, and “organic” traffic stuffed with ads? I’d prefer the private, local, homegrown specialized English-to-bash translator…
      3. lp92 · · focus · HN ↗
        You can run a small model locally with very low latency on consumer hardware not to mention the privacy benefits.
    6. teeskay · · focus · HN ↗
      If it’s one of thing that you want just for English to bash shell commands, I will create AST, it is deterministic, exceptionally fast, no tokens so no need to fine tune existing model, please let me know your thoughts.
      1. brainless · · focus · HN ↗
        You can go quite far using a human language to Bash grammar based setup but at some point the input prompts are harder to translate. The OP has existing projects that work with AST quite deeply so I assume they know about that already.

        I am building a natural language to CSV&#x2F;Excel commands for a &quot;wrangler&quot; type desktop app. Same issues. The MVP is being built with parsers of sorts, entirely code generated. Then I want to fine-tune a tiny model at some point.

        <a href="https:&#x2F;&#x2F;github.com&#x2F;brainless&#x2F;baho" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;brainless&#x2F;baho

    7. okamiueru · · focus · HN ↗
      Golden age before the age that ends humanity. Not talking about any &quot;rogue AI&quot;, just the known statistical models of what is coming due to climate change.
      1. verdverm · · focus · HN ↗
        Do those statistical models account for declining birth rates or are they based on prior population growth projections?
      2. [deleted] · · focus · HN ↗

        [deleted]

    8. xhevahir · · focus · HN ↗
      &gt; what a time to be alive!

      It&#x27;s good to hear you&#x27;re enjoying yourself, but I suggest retiring that expression. It&#x27;s really beginning to grate.

      1. willy_k · · focus · HN ↗
        It’s not so good to hear your pessimism, but I suggest retiring spreading it online. It’s really beginning to grate.
      2. JSR_FDED · · focus · HN ↗
        Ehh, I’m more annoyed by people starting comments with “ehh”
    9. torginus · · focus · HN ↗
      Sorry for the aside, but I noticed half the usecase of AI is fixing the awful DX.
      1. oDot · · focus · HN ↗
        I appreciate the aside. Interesting observation
      2. verdverm · · focus · HN ↗
        I&#x27;m literally working on context&#x2F;harness engineering right now (a set of opencode plugins)

        Aside on the aside, I welcome this new era of really personal software. Not Ai&#x27;s being sycophants, rather being able to easily and quickly change, adapt, or extend software I am not familiar with.

    10. verdverm · · focus · HN ↗
      Seriously, I&#x27;m using a Qwen 3.8 27B on the homelab, distilled from supposed Fable traces. Regardless, the difference is notable, less thinking, better output. Distilled &#x2F; heavy quant is better than the original (imv)

      <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;vwdubb&#x2F;Qwen3.8-27B-Fable-Distill-NVFP4&#x2F;tree&#x2F;main" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;vwdubb&#x2F;Qwen3.8-27B-Fable-Distill-NVFP...

      side quest, are fable distillations only wrong when it&#x27;s another country?

    11. amrrs · · focus · HN ↗
      Did your Astra do any RL or just SFT? did it make up any benchmark to ensure the fine-tuning was a success?
    12. yashthakker · · focus · HN ↗

      [dead]

    13. soundworlds · · focus · HN ↗
      See, you should now share it, so others can benefit without everyone having to do the same re-training :)
    14. ByteOfWood · · focus · HN ↗
      Here&#x27;s a similar project for those who want to replicate: <a href="https:&#x2F;&#x2F;github.com&#x2F;ThorOdinson246&#x2F;whatisit-nl2sh" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;ThorOdinson246&#x2F;whatisit-nl2sh

      Not my project

    15. peab · · focus · HN ↗
      Wow that&#x27;s awesome
    16. brainless · · focus · HN ↗
      I have been trying a mix of fine-tuning and I am amazed that most people do not see this coming.

      A tiny, smaller than 1b parameter model, fine-tuned, can kick ass for constrained work. I do not have a lot of budget, I fine-tune only on a 16GB M4 Mac Mini. But that also tells me the potential is wild. Progress has been slow since I moonlight on this.

      I have been trying to build a set of models + agents for full-stack development, where each model does only a small piece, like take user prompt and break into backend&#x2F;frontend tasks. Then a Rust+Diesel model, a Rust+Auxum model, a Solid+Router model and so on. I know this is wild but this is just theory - can 5 or 6 Qwen 3.5 0.8b models do full-stack web development? My hunch says they can, better than what most people expect. Heck, with a good harness, it might beat all the cheaper models for the specific task, like Haiku or Luna.

      1. RugnirViking · · focus · HN ↗
        remember to test against benchmarks. I would love to hear about your progress.
      2. jordz · · focus · HN ↗
        I think a sub-set of people see this coming, I also think it isn’t just fine tuning open weight LLM models. A few people I know who are thinking along the same lines with architectures like BERT etc.

        That being said it’s much easier at the moment to continue to use the frontier providers for most general tasks, that is the argument I’ve heard.

        For creating these types of fine tuned local models, on constrained hardware for inference, I do think this is the way to go for specific tasks too!

        1. pitched · · focus · HN ↗
          The issue I find with this is that the frontier models still outperform the small finetuned model on its specific task. So much so that the ROI on doing fine tunes is likely negative. I would love to hear some specific example where it did provide value though, if any has any. That would be helpful to start being able to find similar cases.
          1. brainless · · focus · HN ↗
            There are lots of distilled models on huggingface that are much smaller than, say, Opus. They are distilled from Opus or Fable and show clear improvements. I do not have the budget to fine-tune a 30b or more parameter model but from my tiny model experiments, the results are quite clear. Again, I have only a couple small tests.

            Have you actually fine-tuned yourself? Email categorization comes to mind and there are tons of non-LLM approaches even that will give fantastic results. How did spam filters work before LLM?

            I think LLMs just made us think that is the only way. It is not.

        2. williamse · · focus · HN ↗

          [dead]

        3. awwaiid · · focus · HN ↗
          Maybe we&#x27;re getting into a world where our general models can make their own separate fine-tuned models as utilities just like they would a bash&#x2F;python script. If making a fine-tune is cheap and fast then it can also be throw-away and constantly improved.
    17. onel · · focus · HN ↗
      This is a really good story. I love it. Really hope you can share your experience, either blog post or GitHub repo
    18. busfahrer · · focus · HN ↗
      I use this solution for your exact use case:

      I have a single command that fires up llama.cpp on cpu only using gemma4 e2b, answers a single question from the command line and exits. This takes about 3 seconds to load from an SSD, and is smart enough to solve exactly these &quot;remind me of the syntax&quot; scenarios if you dont wanna switch to a browser.

      1. GCUMstlyHarmls · · focus · HN ↗
        I have no experience doing this, and I dont mean this to be snarky: what is the power (and thermal) usage of running your CPU only LLM?

        When I do something &quot;heavy&quot;, my 9800x3D will kick on the fans and start making lot of heat. This is fine when I&#x27;m intentionally doing say a transcode, but if I&#x27;m just querying syntax it could get pretty annoying. Do you &quot;feel&quot; it? I know that will be pretty machine dependent.

    19. everforward · · focus · HN ↗
      I vaguely recall a project from a while back that did something similar without LLMs.

      I’m really pushing my recall, but I want to say it was written in Ruby and stored pre-configured commands that it just did traditional search over.

      I vaguely recall it working okay because 99.99% of the questions people asked were the same (“tar command to gzip a directory and strip the prefix” is something I google like once a month).

    20. ijidak · · focus · HN ↗
      Brilliant.

      Can you (or your agent) please write a tutorial or share some good links.

      (Namely the fine tuning part.)

      I&#x27;d like to learn how to do this as well.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.