‹ BackHN Continuity

Thread

One month coding with GLM 5.3 Flash

228 points · 180 comments · ThibWeb

  1. epistasis · · focus · HN ↗
    One thing about these numbers that's absolutely shocking to me is how low the energy use is:

    > That model’s usage was well within our budget ($68, about 4kWh of energy use / 365 grams of carbon emissions).

    The energy cost is literally 1% of the total cost. For context, 4kWh of energy would drive you about 15 miles in an EV, about half of the average person's daily driving miles. It's boiling 10 gallons of water.

    With the talk of AI data centers' impact on the world, you'd think this would be 10x to 100x the amount of energy in order to get the effects they're using here.

    My takeaway: the AI data center buildout is an overbuild probably at least as large as the fiber buildout that left us with so much dark fiber. If not even bigger. The only thing that will save the economy is the inability of NVIDIA and chip fabs to produce enough chips to match the buildout planned.

    1. scottcha · · focus · HN ↗
      On our journey at Neuralwatt (we are likely the provider he's using as we are the one that does all the energy observability and reporting in our cloud and I'm the CTO there) we quickly discovered that there is a large disconnect between what is in the press (like the rest of the AI narrative very subject to worst case but possibly unlikley future extrapolation of current trends) and what we see on the ground. We actually spend most of our time focused on various way energy constraints (including maximizing tokens/joule while maintaining perf) manifest in current datacenters and provide better ways to get more tokens out of the energy that are already in these data centers or already available but hard to utilize on the grid. So its really a technical constraint problem <it>today</it> rather than an impact problem and I think the press narrative could be better about this.

      Regarding total costs relative to the pure energy costs it is multiple orders of magnitude different but also realize in the datacenter the energy is the pure commodity while almost every other component has huge margins driven by lack of supply. I do think over time this might get closer together (more competition on HW might lower margins) while energy might become more of a bottle neck (raising the energy prices).

      1. polytely · · focus · HN ↗
        Can you say anything about how much variance is there in energy use between models and effort levels. I'm very curious about the difference between the fronteer and things like GLM 5.3 flash.
        1. scottcha · · focus · HN ↗
          We publish live stats comparing model energy here: <a href="https:&#x2F;&#x2F;portal.neuralwatt.com&#x2F;energy-pricing" rel="nofollow">https:&#x2F;&#x2F;portal.neuralwatt.com&#x2F;energy-pricing though since we are ZDR we don&#x27;t really have any information about what effort levels these are run at but I&#x27;ll try and put together some sort of analysis. Its not entirely clear cut as higher effort certainly uses a lot more energy but you might take fewer requests to solve a task.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.