‹ BackHN Continuity

Thread

One month coding with GLM 5.3 Flash

228 points · 180 comments · ThibWeb

  1. epistasis · · focus · HN ↗
    One thing about these numbers that's absolutely shocking to me is how low the energy use is:

    > That model’s usage was well within our budget ($68, about 4kWh of energy use / 365 grams of carbon emissions).

    The energy cost is literally 1% of the total cost. For context, 4kWh of energy would drive you about 15 miles in an EV, about half of the average person's daily driving miles. It's boiling 10 gallons of water.

    With the talk of AI data centers' impact on the world, you'd think this would be 10x to 100x the amount of energy in order to get the effects they're using here.

    My takeaway: the AI data center buildout is an overbuild probably at least as large as the fiber buildout that left us with so much dark fiber. If not even bigger. The only thing that will save the economy is the inability of NVIDIA and chip fabs to produce enough chips to match the buildout planned.

    1. stkdump · · focus · HN ↗
      I have a computer with a 5090 on a smart plug at home running Qwen3.8 27B for agentic coding. On a busy day it can use around 5kWh, though on most days it is around 2kWh. I am sure cloud is more efficient because there is probably efficiency in running many parallel streams, some of the models have fever than 27B active parameters and the power draw of the non-GPU components is also spread over more GPUs. But still I also believe that the energy use numbers you get from the inference providers are a bit "beautified". After all, they still fight a political battle and have to show that it isn't all so bad.

      Having said that, seeing the incredible progress of models throughout this year, I also strongly believe that the planned buildout is overeager. Even I with my gaming hardware often run out of instructions to give. And the smarter the models get that I can run, the less I will be able to saturate my hardware. Is it because of my lack of creativity of which kinds of tasks I can give to AI? Maybe a bit, but currently I can't believe that I am that far away.

      1. rhdunn · · focus · HN ↗
        I suspect that some/a large portion of the build out is due to OpenAI/Anthropic/etc. just throwing hardware at the problem -- why bother trying to make training and inference efficient when you can just throw hardware/money at the problem. I remember NVIDIA talking about their supercomputer clusters with a large number of interconnected GPUs.

        I think the power savings and efficiencies (that the large companies have also benefited from) have come from 2 areas:

        1. open source and local AI enthusiasts -- think things like llama.cpp, quantization, etc.

        2. Chinese labs and other smaller/research companies like Mistral that are using constrained hardware -- see the various advancements in the various models to reduce compute complexity such as mixture of experts [1], sharing key/value data between a group of layers, etc.

        [1] Though the original idea for mixture of experts comes from a 1991 research paper (<a href="https:&#x2F;&#x2F;huggingface.co&#x2F;blog&#x2F;moe" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;blog&#x2F;moe), so maybe a third area is research from Universities, etc.

        1. yorwba · · focus · HN ↗
          The incentives are rather in the opposite direction. Enthusiasts are trying to get models to run at all on the hardware they have, even if the resulting efficiency is poor, whereas a big company spending billions on hardware can afford to hire hundreds of performance engineers to tune their systems end-to-end for maximum efficiency, since the expense pays for itself even if they only manage to eke out a 1% improvement. I.e. why not make training and inference efficient when you can just throw money at the problem?
          1. darkwater · · focus · HN ↗
            Because at the moment they have basically infinite money and it&#x27;s almost always faster to throw hardware at the problem than optimizing the solution, given infinite money.
            1. yorwba · · focus · HN ↗
              Given infinite money, after buying up the finite amount of hardware available, you&#x27;ll still have infinite money left to optimize it for maximum efficiency. And at any finite budget, both the optimal amount to spend on hardware and the optimal amount to spend on efficiency grow with the total budget available, although their relative shares may vary. So while an enthusiast may refuse to invest in hardware at all and focus 100% on efficiency, their efforts will still be outclassed by a large company allocating 99% to hardware and 1% to efficiency.
            2. epistasis · · focus · HN ↗
              That&#x27;s simply untrue, they do not have infinite money nor infinite hardware. They have far stricter limits on hardware than they do on development effort to get more out of their finite hardware.
              1. rhdunn · · focus · HN ↗
                They are acting like they have infinite money (huge annual losses [1], [2], [3]) and infinite hardware (aiming to build a large number of data centres with a lot of GPU hardware [1], [3], [4]) even if that isn&#x27;t the case.

                [1] <a href="https:&#x2F;&#x2F;www.morningstar.com&#x2F;stocks&#x2F;anthropics-leaked-financials-reflect-fast-growth-not-2-trillion-valuation" rel="nofollow">https:&#x2F;&#x2F;www.morningstar.com&#x2F;stocks&#x2F;anthropics-leaked-financi...

                [2] <a href="https:&#x2F;&#x2F;www.wheresyoured.at&#x2F;anthropics-profitability-swindle&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.wheresyoured.at&#x2F;anthropics-profitability-swindle...

                [3] <a href="https:&#x2F;&#x2F;fortune.com&#x2F;2026&#x2F;06&#x2F;16&#x2F;openai-financials-leaked-losses-revenue-profit&#x2F;" rel="nofollow">https:&#x2F;&#x2F;fortune.com&#x2F;2026&#x2F;06&#x2F;16&#x2F;openai-financials-leaked-loss...

                [4] <a href="https:&#x2F;&#x2F;www.forbes.com&#x2F;sites&#x2F;paulocarvao&#x2F;2025&#x2F;12&#x2F;06&#x2F;why-openais-ai-data-center-buildout-faces-a-2026-reality-check&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.forbes.com&#x2F;sites&#x2F;paulocarvao&#x2F;2025&#x2F;12&#x2F;06&#x2F;why-open...

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.