‹ BackHN Continuity

Thread

Ember-1

589 points · 249 comments · gmays

  1. GodelNumbering · · focus · HN ↗
    This is the golden age of model training. Some days ago, I decided I wanted a local CPU only model that can perform exceptionally well for English to Bash translation (to avoid the googling for command syntax). I got a bunch of subagents to generate large amount of training data (140k+ samples), got the Qwen 3 0.6B base model, pointed Astra at it, and off to the races. It trained for 2 days (on and off) and I got a surprisingly good model for my task! The total active time I spent was a few hours. And it is still improving, what a time to be alive!
    1. shriphani · · focus · HN ↗
      what hardware are you using to train?
      1. GodelNumbering · · focus · HN ↗
        I didn't have a local GPU, so I asked it to go out and find hardware. It found a google TPU v6e which seemed reasonably priced. I gave it my google api key. I told it to use TPU only when training and bring it down afterwards. That's about it.
        1. libria · · focus · HN ↗
          > I gave it my google api key

          This is the part where the narrator looks at the camera and says "Don't try this at home, kids!"

          1. MisterMunchkin · · focus · HN ↗
            You’re absolutely right, I shouldn’t have rented a 200 GPU cluster for $35,000/hour. That’s on me.

            [Search: Can I refund Google cloud?]

            It looks like we’re not able to ask for a refund since we did actually use all of that compute intentionally.

            Would you like me to write you a pleading email to send to the support team?

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.