‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

819 points · 362 comments · snehesht

  1. a11r · · focus · HN ↗
    I&#x27;m a little skeptical of going below 4-bit quants due to the potential for significant degradation in quality. I&#x27;m running 4-bit quants on an RTX Pro 6000 rented for approximately $1&#x2F;hour and getting about 1.2 million tokens out and 40 million tokens in per hour with caching. The quality of 4-bit quant is good enough for difficult but well-scoped coding tasks. Here is the inference stack I am using: <a href="https:&#x2F;&#x2F;www.reddit.com&#x2F;r&#x2F;BlackwellPerformance&#x2F;s&#x2F;FrKwk3GoDK" rel="nofollow">https:&#x2F;&#x2F;www.reddit.com&#x2F;r&#x2F;BlackwellPerformance&#x2F;s&#x2F;FrKwk3GoDK
    1. sudo_cowsay · · focus · HN ↗
      For the 1 dollar per hour, may I ask (without trying to incite things) why you would not just buy a ChatGPT Plus subscription or other 20 dollar subscriptions? Do you like the privacy?
      1. mannanj · · focus · HN ↗
        maybe ethics as well. maybe people like sleeping well at night.
        1. tfrancisl · · focus · HN ↗
          [This post has been removed for discussing politics.]
      2. manmal · · focus · HN ↗
        Those plans are super slow right now. Nowhere near 100t&#x2F;s.
        1. martianvoid · · focus · HN ↗
          I think Opus 5.5 is now fast enough with the inference stack improvements that they did. Sonnet 5.5 is around 140 tokens&#x2F;s when I had measured it last time. A $20 subs on claude and using only sonnet 5.5 will last a lot
          1. c16 · · focus · HN ↗
            I was firmly wanting to move towards local llm, and since the Sonnet 5.5 and recent Opus speed increases, I&#x27;m putting that on hold until local inference speeds can be improved. Not sure how much juice there is to squeeze there, but I&#x27;m hopeful.
        2. azath92 · · focus · HN ↗
          While this may be true, their throughput is practically unbounded. this has been a primary challenge to using local models at work, where its normal for me to have 1-6 sessions churning away at once. So while 100t&#x2F;s on a single session is very fast, theres no way you could get a cumulative ~500t&#x2F;s from a consumer setup like this, do the bandwitdth bottleneck on a single card&#x2F;cpu&#x2F;ram&#x2F;ssd (tends to be gpu ram bandwidth is the limiter in my setup, but they all have respective limiters to this kind of throughput)
        3. Almondsetat · · focus · HN ↗
          Does it matter if those 100 tokens per second are way crappier?
      3. rnd0 · · focus · HN ↗
        Access to models is currently being priced below it&#x27;s actual value. Largely because the market hasn&#x27;t settled yet. Once it does the prices to access AI will increase -signifigantly.

        It will be less of an issue for people who have already got the equipment and the practice running locally than it will be for people who have simply been relying on ChatGPT.

        1. jeremyjh · · focus · HN ↗
          Prices on openrouter are not subsidized and are very affordable.
      4. bad_haircut72 · · focus · HN ↗
        People think claude costs $200&#x2F;month, it actually costs $200&#x2F;month and all your labor to train their models, improve their product, and all your best ideas and work that they will feed back into their company. Its very expensive
        1. komatar · · focus · HN ↗
          You can opt out of model training with a checkbox, though I doubt most companies respect this completely. At the very least, I imagine they find loopholes to exploit.

          But on the other hand I&#x27;m giving them my best and worst ideas and the value I get in return outweighs what I&#x27;m contributing to them. I also don&#x27;t have the problems of initial HW cost and maintenance.

          1. radicalbyte · · focus · HN ↗
            &gt; You can opt out of model training with a checkbox

            Do you trust that a company who can&#x27;t even do basic IT ops monitoring are capable of respecting a &quot;no training checkbox&quot;?

          2. wren6991 · · focus · HN ↗
            Great, then they will PII-scrub my sessions before feeding them into the training pool :-)
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.