‹ BackHN Continuity

Thread

Show HN: Fine-tune an 8B model on a 4 GB laptop GPU

135 points · 28 comments · MakazhanAlpamys

  1. kamranjon · · focus · HN ↗
    This seems really interesting - I was curious about this line from the website.

    “The whole post-training stack in one CLI. Soup doctors your data pre-flight, picks the method, writes the config, derives evals from your own data, gates every save, and self-corrects reward hacking mid-run instead of just halting.”

    How does soup auto tune the hyper parameters and make some of these more complex training decisions?

    1. MakazhanAlpamys · · focus · HN ↗

      [dead]

    2. MakazhanAlpamys · · focus · HN ↗

      [dead]

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.