‹ BackHN Continuity

Thread

The AI Race Just Got Awkward

412 points · 463 comments · allisdust

  1. eggbrain · · focus · HN ↗
    Performance optimizations don't just help the western labs, they also help with running more powerful/useful LLMs locally.

    If local LLMs get "good" enough, people will soon paying for subscriptions to ChatGPT and Claude, which hurts their revenue.

    1. kennywinker · · focus · HN ↗
      The only thing preventing this switch from starting in earnest is the data center buildout monopolizing all current and future GPUs
      1. lumost · · focus · HN ↗
        The margins on NVidia datacenter hardware are ... high. At least one order of magnitude larger than a consumer chip.

        Given the recent deepseekv4.1 advances - how good of a 3B model can we make to run on an iphone natively? is it good enough to match common muse/dot use cases for consumers? the phone is already always on.. no need for a cloud server.

        1. christkv · · focus · HN ↗
          any model you run on a phone is not going to do wonders for you battery life
        2. zozbot234 · · focus · HN ↗
          Most phones are not really "always on" in any real sense, a phone on active standby uses very little power and most of it is for its mobile connection. Local AI is best run in a stationary homelab environment, even running it on laptops has its very real problems.
        3. kennywinker · · focus · HN ↗
          There are a few edge models, Spark-X2.5-4B, LFM2.5- in 8b-a1b and 2.6b variants, that are useful for on-device agent-y stuff, but I am doubtful that’s a valuable category. How many iPhone users have never opened Shortcuts in their life? Automating on-phone stuff seems niche.

          I don’t expect the economics of local vs cloud ai to change until either the bubble pops or new ai chips land that can run big models fast with low power demand.

          1. lumost · · focus · HN ↗
            I think the interesting question is whether the average consumer cares? a 4B model should hit roughly the same numbers as GPT5.5 sometime in December/March. That is more than capable of performing a variety of complex work tasks.
            1. kennywinker · · focus · HN ↗
              Skeptical 4B will ever rival GPT5.5, with its estimated 1.5-10T parameters.

              You just can’t compress 1.5T into 4B without losing useful stuff.

              But that doesn’t matter that much. If 4B can compose tools, retrieve info, and not get stuck in loops that is going to handle a lot of use cases

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.