I've been running Qwen 3.8 27b (an opus 4.6 tier model), locally on a 5090 for just over two weeks @ 170 tokens/s. That's a frontier model from 9 months ago running on consumer hardware. Who knows where distillation and pruning gets us in another year.
$9k is a small price to pay to experience the rapturous glory of AGI. I'd easily pay up to 3 times that to comfortably run the superintelligent models released in this post RSI world.
reedf1 · · focus · HN ↗
bix6 · · focus · HN ↗
off_with_their_ · · focus · HN ↗
literalAardvark · · focus · HN ↗