‹ BackHN Continuity

Thread

The AI Race Just Got Awkward

412 points · 463 comments · allisdust

  1. reedf1 · · focus · HN ↗
    I've been running Qwen 3.8 27b (an opus 4.6 tier model), locally on a 5090 for just over two weeks @ 170 tokens/s. That's a frontier model from 9 months ago running on consumer hardware. Who knows where distillation and pruning gets us in another year.
    1. bix6 · · focus · HN ↗
      $9k for a 5090 now? Sheesh.
      1. rubyn00bie · · focus · HN ↗
        In all fairness there are probably a lot of folks who picked one up for around MSRP (even if one of the board partner cards with an MSRP 10-15% over the FE).

        Local inference will have a boom of cheap, powerful, and available cards at some point (even if it isn’t until 2028/2029). At some point the hyperscalers, and frontier labs, will face the capex problems that everyone talks about, and NVidia, AMD, Apple, and Intel will want to keep selling products.

        Powerful, by today’s standard, local inference needs to be accessible to really unlock the “AI” economy long term. It’s just like how the move from mainframes to the PC 40ish years ago unlocked the “computer revolution.”

        1. ethbr1 · · focus · HN ↗
          Especially since a few trends will coincide: memory-optimized model architectures (to save on expensive/rare memory now) + memory glut (because the memory industry, despite its institutional memory, is ramping volume).

          Once hyperscalers stop buying in the quantities they are now, there's going to be a lot of hardware supply to serve by then very hardware efficient models.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.