‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

741 points · 330 comments · snehesht

  1. panny · · focus · HN ↗
    I'm far less interested in how good a big expensive model is on hardware 99% of people can't afford and would rather see what runs best on a chromebook or mobile phone with 8GB of RAM.
    1. liuliu · · focus · HN ↗
      Because that's not possible (to have a GPT 5.6 Sol level model). People won't believe this and will keep dreaming, but intelligence is not free and 8GiB (shared with OS and other processes) is too small to be useful. Whether it is possible for 48GiB or 64GiB (meaning useful for model would be ~16GiB to 24GiB) with external fast storage (SSD), OTOH, is a question mark.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.