‹ BackHN Continuity

Thread

Ember-1

589 points · 249 comments · gmays

  1. GodelNumbering · · focus · HN ↗
    This is the golden age of model training. Some days ago, I decided I wanted a local CPU only model that can perform exceptionally well for English to Bash translation (to avoid the googling for command syntax). I got a bunch of subagents to generate large amount of training data (140k+ samples), got the Qwen 3 0.6B base model, pointed Astra at it, and off to the races. It trained for 2 days (on and off) and I got a surprisingly good model for my task! The total active time I spent was a few hours. And it is still improving, what a time to be alive!
    1. verdverm · · focus · HN ↗
      Seriously, I'm using a Qwen 3.8 27B on the homelab, distilled from supposed Fable traces. Regardless, the difference is notable, less thinking, better output. Distilled / heavy quant is better than the original (imv)

      <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;vwdubb&#x2F;Qwen3.8-27B-Fable-Distill-NVFP4&#x2F;tree&#x2F;main" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;vwdubb&#x2F;Qwen3.8-27B-Fable-Distill-NVFP...

      side quest, are fable distillations only wrong when it&#x27;s another country?

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.