‹ BackHN Continuity

Thread

The AI Race Just Got Awkward

412 points · 463 comments · allisdust

  1. eggbrain · · focus · HN ↗
    Performance optimizations don't just help the western labs, they also help with running more powerful/useful LLMs locally.

    If local LLMs get "good" enough, people will soon paying for subscriptions to ChatGPT and Claude, which hurts their revenue.

    1. kennywinker · · focus · HN ↗
      The only thing preventing this switch from starting in earnest is the data center buildout monopolizing all current and future GPUs
      1. lumost · · focus · HN ↗
        The margins on NVidia datacenter hardware are ... high. At least one order of magnitude larger than a consumer chip.

        Given the recent deepseekv4.1 advances - how good of a 3B model can we make to run on an iphone natively? is it good enough to match common muse/dot use cases for consumers? the phone is already always on.. no need for a cloud server.

        1. kennywinker · · focus · HN ↗
          There are a few edge models, Spark-X2.5-4B, LFM2.5- in 8b-a1b and 2.6b variants, that are useful for on-device agent-y stuff, but I am doubtful that’s a valuable category. How many iPhone users have never opened Shortcuts in their life? Automating on-phone stuff seems niche.

          I don’t expect the economics of local vs cloud ai to change until either the bubble pops or new ai chips land that can run big models fast with low power demand.

          1. lumost · · focus · HN ↗
            I think the interesting question is whether the average consumer cares? a 4B model should hit roughly the same numbers as GPT5.5 sometime in December/March. That is more than capable of performing a variety of complex work tasks.
            1. kennywinker · · focus · HN ↗
              Skeptical 4B will ever rival GPT5.5, with its estimated 1.5-10T parameters.

              You just can’t compress 1.5T into 4B without losing useful stuff.

              But that doesn’t matter that much. If 4B can compose tools, retrieve info, and not get stuck in loops that is going to handle a lot of use cases

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.