‹ BackHN Continuity

Thread

The AI Race Just Got Awkward

412 points · 463 comments · allisdust

  1. eggbrain · · focus · HN ↗
    Performance optimizations don't just help the western labs, they also help with running more powerful/useful LLMs locally.

    If local LLMs get "good" enough, people will soon paying for subscriptions to ChatGPT and Claude, which hurts their revenue.

    1. ericol · · focus · HN ↗
      > people will soon *stop paying

      Think you missed a word there.

    2. kennywinker · · focus · HN ↗
      The only thing preventing this switch from starting in earnest is the data center buildout monopolizing all current and future GPUs
      1. lumost · · focus · HN ↗
        The margins on NVidia datacenter hardware are ... high. At least one order of magnitude larger than a consumer chip.

        Given the recent deepseekv4.1 advances - how good of a 3B model can we make to run on an iphone natively? is it good enough to match common muse/dot use cases for consumers? the phone is already always on.. no need for a cloud server.

        1. christkv · · focus · HN ↗
          any model you run on a phone is not going to do wonders for you battery life
        2. zozbot234 · · focus · HN ↗
          Most phones are not really "always on" in any real sense, a phone on active standby uses very little power and most of it is for its mobile connection. Local AI is best run in a stationary homelab environment, even running it on laptops has its very real problems.
        3. kennywinker · · focus · HN ↗
          There are a few edge models, Spark-X2.5-4B, LFM2.5- in 8b-a1b and 2.6b variants, that are useful for on-device agent-y stuff, but I am doubtful that’s a valuable category. How many iPhone users have never opened Shortcuts in their life? Automating on-phone stuff seems niche.

          I don’t expect the economics of local vs cloud ai to change until either the bubble pops or new ai chips land that can run big models fast with low power demand.

          1. lumost · · focus · HN ↗
            I think the interesting question is whether the average consumer cares? a 4B model should hit roughly the same numbers as GPT5.5 sometime in December/March. That is more than capable of performing a variety of complex work tasks.
            1. kennywinker · · focus · HN ↗
              Skeptical 4B will ever rival GPT5.5, with its estimated 1.5-10T parameters.

              You just can’t compress 1.5T into 4B without losing useful stuff.

              But that doesn’t matter that much. If 4B can compose tools, retrieve info, and not get stuck in loops that is going to handle a lot of use cases

    3. londons_explore · · focus · HN ↗
      For nearly all tasks, I want the fastest and smartest AI model.

      It is vanishingly rare I ask an older model to do any task. Newer bigger and smarter models will just do the task better.

      Therefore, I believe we are nowhere near 'good enough'.

      I never drive my steam engine to work these days. It isn't good enough.

      1. danielmarkbruce · · focus · HN ↗
        I find this too. In fact, recently I've been pushing more and more to the latest and greatest model every time there is an update. It just saves me so much headache.
      2. NoDodgeQuestion · · focus · HN ↗
        I drive my 2018 car to work these days. It is good enough.
      3. eggbrain · · focus · HN ↗
        Right now you are right -- even if I ran my local LLM all day, the quality is not nearly as great, and it runs slowly -- so I use the tier one AI subscription services as they are faster and smarter. But that might only be true for a limited amount of time, and a limited number of circumstances.

        To borrow your steam engine analogy, if local LLMs get as good as a Toyota Prius, even if OpenAI / Anthropic offer Ferraris, most people will be happy with their Prius as their daily driver.

        Similarly, if the big labs start raising prices or cutting usage, you won't be able to use it as much as you want -- whereas a local LLM will run all day every day without costing you any extra money.

        So right now you are right, but who knows how long that will last.

      4. gretch · · focus · HN ↗
        > For nearly all tasks, I want the fastest and smartest AI model.

        This is not nearly true for everyone else in the world.

        For example, think about the world in ~2021 pre-LLM. Would anyone say the sentence "I only want the fastest and smartest humans working on my project"?

        No of course not. Most people don't want to pay $10 million dollar salary to the best programmers in the world. They prefer to pay $200k salary to a median programmer and that's good enough for their ecommerce website.

        1. tyre · · focus · HN ↗
          I mean people said this all the time, but wouldn’t pay for it (as you said), couldn’t recruit for it, and definitely couldn’t retain them.

          But so many teams said they wanted to Raise the Bar to infinity and hire a World Class Team.

      5. settsu · · focus · HN ↗
        Genuine ELI5 question: in this metaphor, which "old" models are steam engines at this point? Which are Model Ts, etc.
      6. monkpit · · focus · HN ↗
        I think a scooter or a bus is a more apt comparison than a steam engine. The scooter and bus can both get the job done with some acceptable trade-offs, depending on your circumstances and what you’re willing to accept.

        However, there are many use cases where they aren’t the right tool for the job.

      7. schmookeeg · · focus · HN ↗
        This will change when out of control billing gets noticed. Then us devs will, I presume, get token rations.

        I can see my new Thursday afternoon "oh chit" moment being that I didn't complete my weekly task because i torched all of those tokens M-W doing task/ticket grooming using the hot hot model instead of the dodo model with jira mcp connector :D

    4. aleqs · · focus · HN ↗
      I would argue performance optimizations help with local/open models but hurt openai and anthropic - because open and or cheap/alternative models are threat to those companies. There is a fundamental contradiction/conflict between the prevelance of open models and the financial success of openai and anthropic. That is why they are doing everything they can to kill any open/cheap/efficient/chinese models (take a look at this thread - it was top of HN 40 mins ago, with very high engagement... now it is buried in page 5... totally normal and legit).
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.