I've been running Qwen 3.8 27b (an opus 4.6 tier model), locally on a 5090 for just over two weeks @ 170 tokens/s. That's a frontier model from 9 months ago running on consumer hardware. Who knows where distillation and pruning gets us in another year.
Indistinguishable or very mildly better. But it's considerably faster. Some portion of that is also probably down to improvements in model harnesses, I've been using opencode.
reedf1 · · focus · HN ↗
oidar · · focus · HN ↗
reedf1 · · focus · HN ↗
jeffrallen · · focus · HN ↗