‹ BackHN Continuity

Thread

One month coding with GLM 5.3 Flash

228 points · 180 comments · ThibWeb

  1. epistasis · · focus · HN ↗
    One thing about these numbers that's absolutely shocking to me is how low the energy use is:

    > That model’s usage was well within our budget ($68, about 4kWh of energy use / 365 grams of carbon emissions).

    The energy cost is literally 1% of the total cost. For context, 4kWh of energy would drive you about 15 miles in an EV, about half of the average person's daily driving miles. It's boiling 10 gallons of water.

    With the talk of AI data centers' impact on the world, you'd think this would be 10x to 100x the amount of energy in order to get the effects they're using here.

    My takeaway: the AI data center buildout is an overbuild probably at least as large as the fiber buildout that left us with so much dark fiber. If not even bigger. The only thing that will save the economy is the inability of NVIDIA and chip fabs to produce enough chips to match the buildout planned.

    1. ThibWeb · · focus · HN ↗
      I agree but it does worry me how fast my usage is increasing. Two months ago I was using 10x less tokens and probably not much more than 5kWh on inference. This month about 30kWh on inference. If it becomes more affordable, is there going to be another jump? Not quite sure
      1. rapatel0 · · focus · HN ↗
        Yeah but performance / watt will decrease - Quadratically with improved chip scaling - orders of magnitude with improved model efficiency
        1. skew-aberration · · focus · HN ↗
          Quadratically? As in, reducing power used for (external) communication/memory busses?
          1. rapatel0 · · focus · HN ↗
            It's a very very very rough approximation, but scaling laws (admittedly weaking over time) yield ~x^2 from a size, ~x^2 from a dynamic power, and more of the data movement is moving from PCB interconnect to on package which looks more similar to chip scaling.

            It's a finger in the wind approximation with lots of confounders, but historically it seems to work out that way for most circuit things.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.