‹ BackHN Continuity

Thread

The AI Race Just Got Awkward

412 points · 463 comments · allisdust

  1. reedf1 · · focus · HN ↗
    I've been running Qwen 3.8 27b (an opus 4.6 tier model), locally on a 5090 for just over two weeks @ 170 tokens/s. That's a frontier model from 9 months ago running on consumer hardware. Who knows where distillation and pruning gets us in another year.
    1. redanddead · · focus · HN ↗
      Well how’s it been so far
      1. off_with_their_ · · focus · HN ↗

        [dead]

    2. oidar · · focus · HN ↗
      What are you thoughts on it's performance compared to 4.6?
      1. reedf1 · · focus · HN ↗
        Indistinguishable or very mildly better. But it's considerably faster. Some portion of that is also probably down to improvements in model harnesses, I've been using opencode.
        1. jeffrallen · · focus · HN ↗
          Yeah, the Shelley agent (from exe.dev) loves Qwen 3.8, they kicked ass on a Django app for me today.
    3. bix6 · · focus · HN ↗
      $9k for a 5090 now? Sheesh.
      1. off_with_their_ · · focus · HN ↗
        $9k is a small price to pay to experience the rapturous glory of AGI. I'd easily pay up to 3 times that to comfortably run the superintelligent models released in this post RSI world.
        1. literalAardvark · · focus · HN ↗
          Except you can do that cheaper by renting compute
      2. bitexploder · · focus · HN ↗
        Well, I have a $750 card that runs at about 50-60% of that token rate :)
        1. iN7h33nD · · focus · HN ↗
          which one?
          1. bitexploder · · focus · HN ↗
            V100S 32GB, I have had Claude optimizing it for about a week and it is already at around 900 t/s prefill, 90-100 t/s output in Pi on coding tasks. There is also a Ninfer fork for the v100 but it requires a custom format. I am working on upstream Unsloth with GGUF 4-bit quant.

            (I also have flash next running even faster on this machine, something a single 5090 can do, with expert cache/pinning, but not quite as fast) :)

      3. rubyn00bie · · focus · HN ↗
        In all fairness there are probably a lot of folks who picked one up for around MSRP (even if one of the board partner cards with an MSRP 10-15% over the FE).

        Local inference will have a boom of cheap, powerful, and available cards at some point (even if it isn’t until 2028/2029). At some point the hyperscalers, and frontier labs, will face the capex problems that everyone talks about, and NVidia, AMD, Apple, and Intel will want to keep selling products.

        Powerful, by today’s standard, local inference needs to be accessible to really unlock the “AI” economy long term. It’s just like how the move from mainframes to the PC 40ish years ago unlocked the “computer revolution.”

        1. ethbr1 · · focus · HN ↗
          Especially since a few trends will coincide: memory-optimized model architectures (to save on expensive/rare memory now) + memory glut (because the memory industry, despite its institutional memory, is ramping volume).

          Once hyperscalers stop buying in the quantities they are now, there's going to be a lot of hardware supply to serve by then very hardware efficient models.

    4. teaearlgraycold · · focus · HN ↗
      Frontier from 9 months ago? I don’t know about that. But it sure punches above its weights.
    5. zdragnar · · focus · HN ↗
      Weird, I kinda gave up on 3.8 as anything other than a planner. I had it try to write some basic unit tests for an admittedly complex bit of code and it ran out of context thinking about the problem and exploring random parts of the code base repeatedly before it even wrote a single line. Toning down the thinking helped some, but then it wasn't much better than qwen coder.
      1. the_lucifer · · focus · HN ↗
        Have you attempted some of the "swift" variants of 3.8? I've heard they're super good in terms of toning down thinking without affecting performance
        1. zdragnar · · focus · HN ↗
          The only swift variant I'm aware of is from ukisai, which uses the swift open license, and is NOT open for commercial use for businesses over $1mil. I'm respecting their choice by not using it for work, which means I'm also not using it for my personal projects in case I forget to switch models when I switch projects.

          My day job doesn't have a dedicated enterprise contract with any of the ai vendors so I might trial it to see if it is worth promoting at the company. Part of me is still holding out hope that qwen 4 dials back the overthinking on its own.

          1. the_lucifer · · focus · HN ↗
            > The only swift variant I'm aware of is from ukisai, which uses the swift open license,

            Ah, I overlooked that, since I have a claude sub at work and all my explorations are purely personal. There's another fast version of 3.8: ThinkingCap[1] by bottlecapai but they have the same $1M restriction from what I can see since it's distributed under a PolyForm Small Business 1.0.0 + BottleCap personal-use grant.

            [1]: <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;bottlecapai&#x2F;ThinkingCap-Qwen3.8-27B" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;bottlecapai&#x2F;ThinkingCap-Qwen3.8-27B

    6. newyankee · · focus · HN ↗
      Do you think this trend can continue ? An Opus5.5 equivalent on a slightly bigger local hardware in under a year ?
      1. an0malous · · focus · HN ↗
        I’m not an AI researcher, but it seems like there’s a ton of waste having a universal model that knows everything when any individuals use case requires like generously 10% of what’s stored in the model. Does it even need to have memorized knowledge stored in the model or could it just look up info and docs like humans do? If all you need is the language and intelligence, I think Opus5.5 equivalent intelligence will run on an iPhone within 5 years.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.