I've been running Qwen 3.8 27b (an opus 4.6 tier model), locally on a 5090 for just over two weeks @ 170 tokens/s. That's a frontier model from 9 months ago running on consumer hardware. Who knows where distillation and pruning gets us in another year.
V100S 32GB, I have had Claude optimizing it for about a week and it is already at around 900 t/s prefill, 90-100 t/s output in Pi on coding tasks. There is also a Ninfer fork for the v100 but it requires a custom format. I am working on upstream Unsloth with GGUF 4-bit quant.
(I also have flash next running even faster on this machine, something a single 5090 can do, with expert cache/pinning, but not quite as fast) :)
reedf1 · · focus · HN ↗
bix6 · · focus · HN ↗
bitexploder · · focus · HN ↗
iN7h33nD · · focus · HN ↗
bitexploder · · focus · HN ↗
(I also have flash next running even faster on this machine, something a single 5090 can do, with expert cache/pinning, but not quite as fast) :)