‹ BackHN Continuity

Thread

One month coding with GLM 5.3 Flash

228 points · 180 comments · ThibWeb

  1. epistasis · · focus · HN ↗
    One thing about these numbers that's absolutely shocking to me is how low the energy use is:

    > That model’s usage was well within our budget ($68, about 4kWh of energy use / 365 grams of carbon emissions).

    The energy cost is literally 1% of the total cost. For context, 4kWh of energy would drive you about 15 miles in an EV, about half of the average person's daily driving miles. It's boiling 10 gallons of water.

    With the talk of AI data centers' impact on the world, you'd think this would be 10x to 100x the amount of energy in order to get the effects they're using here.

    My takeaway: the AI data center buildout is an overbuild probably at least as large as the fiber buildout that left us with so much dark fiber. If not even bigger. The only thing that will save the economy is the inability of NVIDIA and chip fabs to produce enough chips to match the buildout planned.

    1. stkdump · · focus · HN ↗
      I have a computer with a 5090 on a smart plug at home running Qwen3.8 27B for agentic coding. On a busy day it can use around 5kWh, though on most days it is around 2kWh. I am sure cloud is more efficient because there is probably efficiency in running many parallel streams, some of the models have fever than 27B active parameters and the power draw of the non-GPU components is also spread over more GPUs. But still I also believe that the energy use numbers you get from the inference providers are a bit "beautified". After all, they still fight a political battle and have to show that it isn't all so bad.

      Having said that, seeing the incredible progress of models throughout this year, I also strongly believe that the planned buildout is overeager. Even I with my gaming hardware often run out of instructions to give. And the smarter the models get that I can run, the less I will be able to saturate my hardware. Is it because of my lack of creativity of which kinds of tasks I can give to AI? Maybe a bit, but currently I can't believe that I am that far away.

      1. Tuna-Fish · · focus · HN ↗
        Batching is extremely powerful for optimizing power use. Your card is spending the vast majority of it's energy use fetching weights from it's memory, and it only gets to use each weight once. Running that same model on a bigger machine with lots of batching gets to amortize the cost of loading a weight across all the parallel requests. N=128 is literally about 50 times more energy-efficient than N=1.
        1. stkdump · · focus · HN ↗
          I don't know. It would surprise me. For long agentic work you shouldn't need that many streams for the KV cache to become larger than the (active) parameters. At 128 streams the model weights would be neglectable. I am using a 22GB model file, i.e. I have 10GB for context, which isn't even the max that the model supports.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.