‹ BackHN Continuity

Thread

One month coding with GLM 5.3 Flash

228 points · 180 comments · ThibWeb

  1. epistasis · · focus · HN ↗
    One thing about these numbers that's absolutely shocking to me is how low the energy use is:

    > That model’s usage was well within our budget ($68, about 4kWh of energy use / 365 grams of carbon emissions).

    The energy cost is literally 1% of the total cost. For context, 4kWh of energy would drive you about 15 miles in an EV, about half of the average person's daily driving miles. It's boiling 10 gallons of water.

    With the talk of AI data centers' impact on the world, you'd think this would be 10x to 100x the amount of energy in order to get the effects they're using here.

    My takeaway: the AI data center buildout is an overbuild probably at least as large as the fiber buildout that left us with so much dark fiber. If not even bigger. The only thing that will save the economy is the inability of NVIDIA and chip fabs to produce enough chips to match the buildout planned.

    1. etdznots · · focus · HN ↗
      Inference needs very little compute, it’s mostly IO. hence the low power draw, training is compute-heavy and uses lots of power
      1. apitman · · focus · HN ↗
        Honest question: if that's the case why do my GPUs draw their full TDP during inference?
        1. etdznots · · focus · HN ↗
          I am not an expert on why, but my vague answer is you are still running near max clock speeds, and you can significantly downclock your gpu and&#x2F;or set power limits on your GPU without seeing much loss in performance, e.g. you can cap a 3090 to 60% of it’s normal power draw and it will lose like 5% of tg performance and 10% of pp performance (<a href="https:&#x2F;&#x2F;www.reddit.com&#x2F;r&#x2F;LocalLLaMA&#x2F;comments&#x2F;1hg6qrd&#x2F;relative_performance_in_llamacpp_when_adjusting&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.reddit.com&#x2F;r&#x2F;LocalLLaMA&#x2F;comments&#x2F;1hg6qrd&#x2F;relativ...)

          I think theoretically this is still very wasteful with lots of compute that’s getting powered and sitting unused but you are limited either by the firmware or by the GPU architecture from pushing power usage even lower without cratering performance. (no one anticipated the demand for relatively low-compute devices with lots of super fast memory)

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.