‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

826 points · 365 comments · snehesht

  1. a11r · · focus · HN ↗
    I&#x27;m a little skeptical of going below 4-bit quants due to the potential for significant degradation in quality. I&#x27;m running 4-bit quants on an RTX Pro 6000 rented for approximately $1&#x2F;hour and getting about 1.2 million tokens out and 40 million tokens in per hour with caching. The quality of 4-bit quant is good enough for difficult but well-scoped coding tasks. Here is the inference stack I am using: <a href="https:&#x2F;&#x2F;www.reddit.com&#x2F;r&#x2F;BlackwellPerformance&#x2F;s&#x2F;FrKwk3GoDK" rel="nofollow">https:&#x2F;&#x2F;www.reddit.com&#x2F;r&#x2F;BlackwellPerformance&#x2F;s&#x2F;FrKwk3GoDK
    1. sudo_cowsay · · focus · HN ↗
      For the 1 dollar per hour, may I ask (without trying to incite things) why you would not just buy a ChatGPT Plus subscription or other 20 dollar subscriptions? Do you like the privacy?
      1. rnd0 · · focus · HN ↗
        Access to models is currently being priced below it&#x27;s actual value. Largely because the market hasn&#x27;t settled yet. Once it does the prices to access AI will increase -signifigantly.

        It will be less of an issue for people who have already got the equipment and the practice running locally than it will be for people who have simply been relying on ChatGPT.

        1. jeremyjh · · focus · HN ↗
          Prices on openrouter are not subsidized and are very affordable.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.