‹ BackHN Continuity

Thread

The AI Race Just Got Awkward

412 points · 463 comments · allisdust

  1. reedf1 · · focus · HN ↗
    I've been running Qwen 3.8 27b (an opus 4.6 tier model), locally on a 5090 for just over two weeks @ 170 tokens/s. That's a frontier model from 9 months ago running on consumer hardware. Who knows where distillation and pruning gets us in another year.
    1. bix6 · · focus · HN ↗
      $9k for a 5090 now? Sheesh.
      1. bitexploder · · focus · HN ↗
        Well, I have a $750 card that runs at about 50-60% of that token rate :)
        1. iN7h33nD · · focus · HN ↗
          which one?
          1. bitexploder · · focus · HN ↗
            V100S 32GB, I have had Claude optimizing it for about a week and it is already at around 900 t/s prefill, 90-100 t/s output in Pi on coding tasks. There is also a Ninfer fork for the v100 but it requires a custom format. I am working on upstream Unsloth with GGUF 4-bit quant.

            (I also have flash next running even faster on this machine, something a single 5090 can do, with expert cache/pinning, but not quite as fast) :)

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.