‹ BackHN Continuity

Thread

One month coding with GLM 5.3 Flash

228 points · 180 comments · ThibWeb

  1. epistasis · · focus · HN ↗
    One thing about these numbers that's absolutely shocking to me is how low the energy use is:

    > That model’s usage was well within our budget ($68, about 4kWh of energy use / 365 grams of carbon emissions).

    The energy cost is literally 1% of the total cost. For context, 4kWh of energy would drive you about 15 miles in an EV, about half of the average person's daily driving miles. It's boiling 10 gallons of water.

    With the talk of AI data centers' impact on the world, you'd think this would be 10x to 100x the amount of energy in order to get the effects they're using here.

    My takeaway: the AI data center buildout is an overbuild probably at least as large as the fiber buildout that left us with so much dark fiber. If not even bigger. The only thing that will save the economy is the inability of NVIDIA and chip fabs to produce enough chips to match the buildout planned.

    1. ThibWeb · · focus · HN ↗
      I agree but it does worry me how fast my usage is increasing. Two months ago I was using 10x less tokens and probably not much more than 5kWh on inference. This month about 30kWh on inference. If it becomes more affordable, is there going to be another jump? Not quite sure
      1. CuriouslyC · · focus · HN ↗
        At some point the agents will be good enough that you can tell them "here's $100, go make me money" and they will, maybe not a lot and not all the time, but the EV given the cost of inference will be positive.
        1. pphysch · · focus · HN ↗
          Good enough at what, fraud? Robotic Ponzi schemes would be an easy way to generate "income".
          1. CuriouslyC · · focus · HN ↗
            More likely they bot will do research on how to make make money and do experiments given the resources available to it.
            1. lukan · · focus · HN ↗
              I guess the question boils down to, what remote work could a human do to earn money over the internet?

              So how many companies/individuals pay a random spam bot contacting them 100€ to upgrade their website? I guess start with targeting gullible and trying to scam them will be the more successful experiment for them. (If run unrestricted)

        2. mannanj · · focus · HN ↗
          Those rewards seem like they would be captured by the providers and AI companies though who would use them first to make themselves money. You would be left with whatever they didn’t pursue with their first movers advantage.
        3. MisterMunchkin · · focus · HN ↗
          If I’m the AI company why would I let you do that, when I could just do it myself and get all of the money?
          1. CuriouslyC · · focus · HN ↗
            Regulation. If it got proven to the point that it was scaled out en mass, it would get hit hard and fast.
            1. monkpit · · focus · HN ↗
              ???

              Then why would any AI company exist when they could just use their money to buy tokens from another AI company and make money for zero effort? There would be no incentive to be a provider.

              Not to mention inflation would grow to match or outpace the rate you could earn on these guaranteed AI gains.

        4. rhyperior · · focus · HN ↗
          If that does become possible then the market will naturally establish efficiency again and the opportunity will disappear.
          1. cyanydeez · · focus · HN ↗
            it'll basically be scam vs scam if that's every what happens.
          2. gunalx · · focus · HN ↗
            If the market would be fully efficient how would it be possible to make money in the market? I font think the market is that fast at rebalancing.
          3. epistasis · · focus · HN ↗
            This seems somewhat obviously true, to some degree, but it all depends on the prompt, I think?

            In the equilibrium, profits fall to zero in capitalism, but there's never equilibrium, everything is changing. Profit comes from insights that others haven't seen, from access to opportunities that others don't have, or through monopolistic control generating economic rents.

            For that AI to make money, it has to have a harness that gives it one of the above things, which does seem possible.

            1. swiftcoder · · focus · HN ↗
              > which does seem possible

              But strictly time-limited, because as soon as someone else figures out what you are doing, there's no moat to prevent them from telling their AI to do it too.

              Markets reach equilibrium very fast when everyone has access to the exact same tools and information.

        5. a3dds · · focus · HN ↗
          Bro.. your brain is fried.

          Take a vacation.

      2. rapatel0 · · focus · HN ↗
        Yeah but performance / watt will decrease - Quadratically with improved chip scaling - orders of magnitude with improved model efficiency
        1. skew-aberration · · focus · HN ↗
          Quadratically? As in, reducing power used for (external) communication/memory busses?
          1. rapatel0 · · focus · HN ↗
            It's a very very very rough approximation, but scaling laws (admittedly weaking over time) yield ~x^2 from a size, ~x^2 from a dynamic power, and more of the data movement is moving from PCB interconnect to on package which looks more similar to chip scaling.

            It's a finger in the wind approximation with lots of confounders, but historically it seems to work out that way for most circuit things.

        2. 233mhz · · focus · HN ↗
          And usage will grow accordingly, as usual.

          My car engine is a miracle of efficiency compared to the stuff we had a hundred years ago, but overall cars pollute more than back then

      3. driverdan · · focus · HN ↗
        To put that in perspective my house uses an average of 1300kWh per month. Adding 30kWh per month would be an increase of 2.3% at a cost of around $4.20, not enough to notice.
        1. spiffytech · · focus · HN ↗
          Neuralwatt prices proportional to the energy used (to cover all the non-energy costs). 30 kWh is $200 of service.
          1. driverdan · · focus · HN ↗
            Sure, they have to cover their other costs and make money. I was trying to put the energy use in perspective.
        2. ThibWeb · · focus · HN ↗
          for me it’s more an environmental consideration than financial? Passivhaus Standard for houses is about 50kWh/m²/year, so 30kWh is getting close to 10% of total
        3. oblio · · focus · HN ↗
          Do you live in the US? You should know in that case that the US is already one of the most wasteful energy users on the planet, far ahead of most other developed countries, even.

          So potentially increasing energy usage even more, especially with likely token usage acceleration, is hardly reason for celebration.

          1. driverdan · · focus · HN ↗
            No one is celebrating and using LLMs isn't necessarily a waste of energy. Where I live isn't relevant, LLMs use the same amount of power regardless of where their users are. You also know nothing about how I use electricity in my home, how I monitor my overall energy use and carbon output, and how I offset it.
            1. nicoburns · · focus · HN ↗
              > You also know nothing about how I use electricity in my home

              I don't

              > Where I live isn't relevant,

              But where you live does matter. Average monthly household energy usage in the EU is 200-400 kWh. In the US it's more like 800-900kWh. So what you consider "normal energy usage" may vary considerably by where you live.

              At 200-400kWh total usage, 30kwh starts to look a lot more significant. So you should at least consider the possibility that it's not that energy usage from AI is insignificant, but that your existing energy usage is wasteful.

            2. oblio · · focus · HN ↗
              You're American and the same 2.3% of energy you use would probably represents 5-10% of what the average Japanese person uses, and likely 15-20% of what the average Indian or Ethiopian uses.

              My point isn't to attack you personally (I'd rather attack your country's disastrous energy policies) but to point out that LLMs will constitute a huge increase in our energy consumption at a time where we haven't transitioned to renewable energy and we're probably 20-30 away from it.

        4. [deleted] · · focus · HN ↗

          [deleted]

        5. 233mhz · · focus · HN ↗
          > my house uses an average of 1300kWh per month.

          And the average European house uses like 3000kwh.... per year. We're already wasting insane amount of energy, adding 2-10% more is definitely huge

    2. killingtime74 · · focus · HN ↗
      Your only talking about variable direct energy. Does it take into account the entire lifecycle, building the data centre, running the cooling, building the chips, the % the chips are not utilised.
      1. ThibWeb · · focus · HN ↗
        No those numbers are GPU only. See <a href="https:&#x2F;&#x2F;cleerdash.sustainableaigroup.com&#x2F;" rel="nofollow">https:&#x2F;&#x2F;cleerdash.sustainableaigroup.com&#x2F; for a fuller model
      2. epistasis · · focus · HN ↗
        &gt; Does it take into account the entire lifecycle, building the data centre

        My entire point is that 99% of the dollar cost of running these models goes to things other than the GPU power. The capex cost to building cost to GPU cost to storage&#x2F;networking&#x2F;chasses&#x2F;wiring plus the other operation costs dwarf the electricity. Even the other electricity costs, lets say double it for all the supporting compute, plus another 25% for a 1.25 PUE, and you&#x27;re at 2.5% of all-in cost of running these models is from electricity.

        The non-electricity costs are massive and the constraints on fabs, etc. will drive the amount of the AI build far more than energy availability.

    3. smartmic · · focus · HN ↗
      Related to this: see the difference between China and US here: <a href="https:&#x2F;&#x2F;aidatacenterindex.com&#x2F;compare&#x2F;united-states-vs-china&#x2F;" rel="nofollow">https:&#x2F;&#x2F;aidatacenterindex.com&#x2F;compare&#x2F;united-states-vs-china...

      I have the feeling China is somehow ahead when it comes to energy (and cost) efficiency for AI usage. After all, the two are in a direct competition, and this difference is significant. Or is the &quot;hyper&quot; scaling of energy hungry datacenters in US part of a bubble?

    4. eli · · focus · HN ↗
      It would be a pretty big deal if the average person&#x27;s car became 50% less energy efficient, no?
    5. jjcm · · focus · HN ↗
      &gt; AI datacenter buildout is an overbuild

      a.) we’re supply constrained

      b.) only 3% of households pay for AI

      Inference amounts will continue to grow heavily.

    6. timmmmmmay · · focus · HN ↗
      &quot;the talk of AI data centers&#x27; impact on the world&quot; has been wildly exaggerated and you can see here that this is the least impact of any major new technology in the history of industrialization
      1. DragonStrength · · focus · HN ↗
        Exaggerated or poorly represented? One candidate for governor in my state likens it to the railroad bubble. Satya Nadella makes the same comparison. What happens when the bubble pops? We&#x27;re seeing data centers outbid industrial plants for sites. Plants which would provide 100x as many jobs as a data center and not be abandoned in 2028 when the chips still aren&#x27;t available and contracts come due.

        Oh right, just some rubes in flyover states that don&#x27;t know what&#x27;s good for them.

    7. api · · focus · HN ↗
      Yeah the whole data center environmental panic seemed astroturfed to me. I took a look at the numbers and it’s not that bad. If you telework one day instead of commuting you make up for over a week of heavy AI use, and the water use is on par with an average golf course.

      There are noise issues in some places. But the panic is excessive. Like nuts.

      Maybe environmental panics are to the left what moral panics about stuff like trans people are to the right.

      1. epistasis · · focus · HN ↗
        Don&#x27;t need any astroturfing to explain the data center panic, it&#x27;s well within organic NIMBYism. I&#x27;ve been trying to get housing built in my town for a decade, the crap stuff say about housing and the fervor with which they block housing is even greater than the data center panic.

        Add in 1) all job destruction that Sam Altman and other leaders tout, 2) Musk operating an environmental disaster in the most offensive way possible for xAI, and 3) general fear bait a new technology that has such uncertain consequences, and I&#x27;m surprised the backlash isn&#x27;t stronger.

        Go talk to community members about any change in the buildings in their area and the data center backlash fits right in line with normal responses.

      2. duskdozer · · focus · HN ↗
        How is it excessive and nuts? If anywhere else is like it is by me, they&#x27;ve been ramming a bunch of giant warehouses at the least, and the increased noise both from the operation itself and the loss of natural barriers blocking other noise, light, and air pollution is really noticeably degrading things. And from what I understand, data centers are even worse due to their increased use of cooling and things like gas generators. It shouldn&#x27;t come as a surprise to anyone that people are pissed off and increasingly so.
    8. cyanydeez · · focus · HN ↗
      I&#x27;m running Qwen3.8-Flash-Next; other than speed on a 395+, it does most things that are properly planned out.

      I can&#x27;t believe the TAM requires more than it at a 2x speed up. OpenAI, Anthropic built models whose only purpose now is things like research and defense.

      And by defense, obviously, in America, we mean war. killing, etc.

    9. ericd · · focus · HN ↗
      It&#x27;s running a normal central home heat pump&#x2F;air conditioner for one hour, or eight hours of playing on a gaming PC. Yeah, the propaganda around datacenter resource usage has massively outrun the truth, which makes me think there are some very interested parties pushing behind the scenes.
      1. cyberrock · · focus · HN ↗
        I don&#x27;t discount the possibility, but I think part of it is that no one was interested in knowing how many resources normal computing costs. Any one of us who&#x27;s worked on a video platform or cloud storage providers already knew how the sausage was made, but as long as it provided unlimited videos and photo storage, the public didn&#x27;t care.

        Though I do suspect that golf course owners in Arizona and alfalfa farmers in California must be ecstatic.

    10. zozbot234 · · focus · HN ↗
      The AI datacenter buildout is in the gigawatts. A single 1.21 GW data center multiplies that single-user 5.5 W average load (4 kWh per month) by as much as ~200 million. Obviously a widely distributed load is much less impactful on the surrounding environment.

      BTW, according to the article the full GLM 5.3 model is in the &quot;10x to 100x&quot; ballpark. So this is not outside the realm of plausibility.

      1. epistasis · · focus · HN ↗
        Average US electrical grid load is 500GW.

        As we electrify and decarbonize, we are looking at going to 1000-1500GW average load (clean energy sources are 2x-6x more energy efficient than fossil fuels at delivering energy services , so looking at the primary energy flows here you&#x27;ll not only see elimination of the &quot;rejected energy&quot; category, but heating by fossil fuels gets counted as ~100% effect when heat pumps are 200%-600% effect on those terms <a href="https:&#x2F;&#x2F;flowcharts.llnl.gov&#x2F;" rel="nofollow">https:&#x2F;&#x2F;flowcharts.llnl.gov&#x2F; )

        Going to 50GW of AI would be an amazing jump for which we don&#x27;t have fab capacity anytime.

    11. gchamonlive · · focus · HN ↗
      &gt; With the talk of AI data centers&#x27; impact on the world, you&#x27;d think this would be 10x to 100x the amount of energy in order to get the effects they&#x27;re using here.

      This was going so well until this. Everything at scale has environmental impact because you centralize the downside and distribute the upside. This is an important alienation, but it hides the amount of heat, noise and impact on distribution that datacenters have on local infrastructure and environment.

      1. hedora · · focus · HN ↗
        Yes, but, I can run qwen 3.8 flash next (&gt; 100b parameters; ~ opus 4.6) on a 200W tdp desktop.

        So, assume 8 of those, and it’s a space heater per house. I have never been disturbed by my neighbor’s space heater.

        The problem is centralization, not the absolute energy usage. (And also that LLMs are trending to 100x more efficient than the data center sizing assumed).

        1. gchamonlive · · focus · HN ↗
          Decentralization can also cause havoc because of back-EMF for instance, in any case the public infrastructure needs time to adapt to demand.
        2. epistasis · · focus · HN ↗
          Centralization is vastly more energy efficient than decentralization, by orders of magnitude. LLMS love batching, to an insane degree. Running a single conversation through a GPU is about the same cost as doing ~100 in parallel.

          Centralization is a huge huge environmental win, far far more than even cloud computing was compared to tons of inefficient, under utilized racks spread though our office buildings.

          1. z0r · · focus · HN ↗
            Huge environmental win if considering a fixed supply &amp; demand, but Jevons Paradox kicks in with tokens in the cloud
            1. rpdillon · · focus · HN ↗
              This is very true. I&#x27;m more careful with what I ask my local model, since it&#x27;s about 1&#x2F;3 as fast as the cloud models I&#x27;m used to.
          2. hypfer · · focus · HN ↗
            &gt; Centralization is a huge huge environmental win

            In a vacuum, yes. Assuming everyone is a good actor.

          3. rpdillon · · focus · HN ↗
            Totally true. But then it introduces a much greater degree of control, and it creates the aforementioned local concentration of energy usage, noise, etc. So it&#x27;s a tradeoff, more than a strict win.
      2. epistasis · · focus · HN ↗
        What is your point?

        Calculate the scale. Add it up. Do it.

        Look at the numbers and come back to me.

        1. ls612 · · focus · HN ↗
          He is a Marxist (the alienation rhetoric gives it away) he is primarily interested in causing social revolution not in environmentalism or anything else.
          1. [deleted] · · focus · HN ↗

            [deleted]

          2. gchamonlive · · focus · HN ↗
            Labels are nice to fit what you don&#x27;t understand in terms you feel like you understand. The local impact is real. Alienation is important. You are stuck in late 1800s rethoric.
            1. Fidelix · · focus · HN ↗
              You didn&#x27;t deny anything he said
              1. gchamonlive · · focus · HN ↗
                If you think long and hard about it you&#x27;ll see I did.

                But if you must you can call me a moderately progressive Marxist.

                1. simianwords · · focus · HN ↗
                  not beating the allegations
                  1. [deleted] · · focus · HN ↗

                    [deleted]

                  2. 233mhz · · focus · HN ↗
                    This is a techie shit posting forum not a court of law, calm your tits down
        2. gchamonlive · · focus · HN ↗
          My point is the impact is real, measurable, but nobody will do it because nobody can stop the hype train lest they be labeled opponents to the progress
          1. epistasis · · focus · HN ↗
            These days even the AI companies are urging a slowdown if pace.

            If somebody is afraid of speaking their mind, because the are afraid of being labelled an opponent to progress, that&#x27;s some personal issue to work through.

            It is popular and encouraged to be skeptical of AI, there no social opprobrium about it, unless you&#x27;re in an extremist political cult, in which case you probably call it SI.

            1. gchamonlive · · focus · HN ↗
              &gt; there no social opprobrium about it

              You can&#x27;t really say this categorically, because that&#x27;s not a universal experience.

              Maybe you are one of the lucky ones not having immense pressure at work to adopt AI at all costs, but the reality of it is that if I open my mike in a &quot;debate&quot; where I work with all devs and managers, to speak in favor of taking things easy and to give time for people to learn these tools properly, I&#x27;d be frowned upon.

              So while it&#x27;s true that&#x27;s encouraged to be skeptical about the tech, this really depends on the context and saying it doesn&#x27;t clashes hard with my experience of the world.

    12. redanddead · · focus · HN ↗
      &gt; My takeaway: the AI data center buildout is an overbuild probably at least as large as the fiber buildout that left us with so much dark fiber. If not even bigger. The only thing that will save the economy is the inability of NVIDIA and chip fabs to produce enough chips to match the buildout planned.

      I never understand the mindset that leads to these massive leaps in reasoning. I was recruiting a guy, he was a VC GP, he says the same thing. I think he’s wrong. The more we use the more we want to use.

      Do we seriously think our species will jinx the Kardashev energy requirements

      1. epistasis · · focus · HN ↗
        The unjustified leap here seems to be talking about Kardashev every requirements with relation to data centers.

        Every massive technological leap that have big private buildouts results in massive overbuild. It&#x27;s the nature of FOMO when a big civilization changing tech gets introduced.

        Will AI change everything? Are too many data centers planned? The answer is almost certainly yes to both.

        1. redanddead · · focus · HN ↗
          Regardless, we&#x27;ll need the processing power and the energy

          Is that best spent on LLMs, world models, physical AI, who knows. But we&#x27;ll need the infrastructure, what&#x27;s being built is conservative

        2. rpdillon · · focus · HN ↗
          Didn&#x27;t the overbuild of fiber directly lead to Google buying up dark fiber, and then lighting it up later on to power YouTube? I ran across this when I was trying to figure out why YouTube has no real competitors, and the explanations I ran across were basically “Google has an immense economic advantage because of their control over all this extra bandwidth capacity&quot;. I have not further validated that in the couple of years since I read it, but it does make me wonder what excess AI capacity could lead to.
          1. swiftcoder · · focus · HN ↗
            &gt; Didn&#x27;t the overbuild of fiber directly lead to Google buying up dark fiber, and then lighting it up later

            Yes, but the lag there is decade-long, and a bunch of heavily fibre-invested companies went bust before anyone saw the upside.

            In the same vein, it&#x27;s very possible that early movers-and-shakers in the LLM space will overextend and end up bankrupt before LLMs find a long-term profitable niche

            1. rpdillon · · focus · HN ↗
              Yep, the decade-long timeline about matches what I recall. Agree on the culling that&#x27;s coming in the AI space as well. Temporal&#x27;s comment from a couple of days ago about inside vs. outside AI has been thought-provoking as I think about who is going to go bust...
    13. scottcha · · focus · HN ↗
      On our journey at Neuralwatt (we are likely the provider he&#x27;s using as we are the one that does all the energy observability and reporting in our cloud and I&#x27;m the CTO there) we quickly discovered that there is a large disconnect between what is in the press (like the rest of the AI narrative very subject to worst case but possibly unlikley future extrapolation of current trends) and what we see on the ground. We actually spend most of our time focused on various way energy constraints (including maximizing tokens&#x2F;joule while maintaining perf) manifest in current datacenters and provide better ways to get more tokens out of the energy that are already in these data centers or already available but hard to utilize on the grid. So its really a technical constraint problem &lt;it&gt;today&lt;&#x2F;it&gt; rather than an impact problem and I think the press narrative could be better about this.

      Regarding total costs relative to the pure energy costs it is multiple orders of magnitude different but also realize in the datacenter the energy is the pure commodity while almost every other component has huge margins driven by lack of supply. I do think over time this might get closer together (more competition on HW might lower margins) while energy might become more of a bottle neck (raising the energy prices).

      1. [deleted] · · focus · HN ↗

        [deleted]

      2. polytely · · focus · HN ↗
        Can you say anything about how much variance is there in energy use between models and effort levels. I&#x27;m very curious about the difference between the fronteer and things like GLM 5.3 flash.
        1. scottcha · · focus · HN ↗
          We publish live stats comparing model energy here: <a href="https:&#x2F;&#x2F;portal.neuralwatt.com&#x2F;energy-pricing" rel="nofollow">https:&#x2F;&#x2F;portal.neuralwatt.com&#x2F;energy-pricing though since we are ZDR we don&#x27;t really have any information about what effort levels these are run at but I&#x27;ll try and put together some sort of analysis. Its not entirely clear cut as higher effort certainly uses a lot more energy but you might take fewer requests to solve a task.
        2. lhl · · focus · HN ↗
          On a blend of about 300 agentic&#x2F;coding tasks, for GLM-5.3 Flash, low was about 1.3M tok&#x2F;ans, medium was ~2M tok&#x2F;ans, and xhigh&#x2F;max was 4.8M tok&#x2F;ans. This was dominated by output tokens btw (only about a 2X input token spread, but closers to a 5X spread for output tokens from low to xhigh).

          These differ by model and set of tasks, but for GLM-5.3 Flash, the best $&#x2F;pass was medium effort. Max did get about 5% higher scores, but at &gt;2X the $&#x2F;pass.

      3. ThibWeb · · focus · HN ↗
        (OP) yes all energy &#x2F; carbon footprint numbers are from Neuralwatt. ty for your work! Switched from Standard to Pro last week, it’s great
        1. scottcha · · focus · HN ↗
          Awesome, glad it worked well. Amazing writeups!
    14. teaearlgraycold · · focus · HN ↗
      &gt; With the talk of AI data centers&#x27; impact on the world, you&#x27;d think this would be 10x to 100x the amount of energy in order to get the effects they&#x27;re using here.

      Well there&#x27;s also training to consider. But yeah, most of what I see online about datacenters is nonsense. People worried about water use as a top concern are misinformed.

      Look, I absolutely get it if you&#x27;re worried your boss is just waiting for the day to replace you with an LLM. Fucking get organized with your fellow laborers instead of getting distracted by datacenter water use. Call your politicians to rein in billionaires. If anything the ownership class is probably directing online discourse towards electricity and water user to keep people away from class-focused debates.

      1. etdznots · · focus · HN ↗
        No, because it’s a viable strategy that has teeth because of existing regs and agencies, not the battle to fight until the end of the earth, but it’s a pretty big hammer
        1. teaearlgraycold · · focus · HN ↗
          Sure, and I love to see people actually giving a shit and getting engaged locally. There are many things I see - tweets and video from protests - that make me think people should take that energy to more direct issues.
    15. stkdump · · focus · HN ↗
      I have a computer with a 5090 on a smart plug at home running Qwen3.8 27B for agentic coding. On a busy day it can use around 5kWh, though on most days it is around 2kWh. I am sure cloud is more efficient because there is probably efficiency in running many parallel streams, some of the models have fever than 27B active parameters and the power draw of the non-GPU components is also spread over more GPUs. But still I also believe that the energy use numbers you get from the inference providers are a bit &quot;beautified&quot;. After all, they still fight a political battle and have to show that it isn&#x27;t all so bad.

      Having said that, seeing the incredible progress of models throughout this year, I also strongly believe that the planned buildout is overeager. Even I with my gaming hardware often run out of instructions to give. And the smarter the models get that I can run, the less I will be able to saturate my hardware. Is it because of my lack of creativity of which kinds of tasks I can give to AI? Maybe a bit, but currently I can&#x27;t believe that I am that far away.

      1. rhdunn · · focus · HN ↗
        I suspect that some&#x2F;a large portion of the build out is due to OpenAI&#x2F;Anthropic&#x2F;etc. just throwing hardware at the problem -- why bother trying to make training and inference efficient when you can just throw hardware&#x2F;money at the problem. I remember NVIDIA talking about their supercomputer clusters with a large number of interconnected GPUs.

        I think the power savings and efficiencies (that the large companies have also benefited from) have come from 2 areas:

        1. open source and local AI enthusiasts -- think things like llama.cpp, quantization, etc.

        2. Chinese labs and other smaller&#x2F;research companies like Mistral that are using constrained hardware -- see the various advancements in the various models to reduce compute complexity such as mixture of experts [1], sharing key&#x2F;value data between a group of layers, etc.

        [1] Though the original idea for mixture of experts comes from a 1991 research paper (<a href="https:&#x2F;&#x2F;huggingface.co&#x2F;blog&#x2F;moe" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;blog&#x2F;moe), so maybe a third area is research from Universities, etc.

        1. yorwba · · focus · HN ↗
          The incentives are rather in the opposite direction. Enthusiasts are trying to get models to run at all on the hardware they have, even if the resulting efficiency is poor, whereas a big company spending billions on hardware can afford to hire hundreds of performance engineers to tune their systems end-to-end for maximum efficiency, since the expense pays for itself even if they only manage to eke out a 1% improvement. I.e. why not make training and inference efficient when you can just throw money at the problem?
          1. darkwater · · focus · HN ↗
            Because at the moment they have basically infinite money and it&#x27;s almost always faster to throw hardware at the problem than optimizing the solution, given infinite money.
            1. yorwba · · focus · HN ↗
              Given infinite money, after buying up the finite amount of hardware available, you&#x27;ll still have infinite money left to optimize it for maximum efficiency. And at any finite budget, both the optimal amount to spend on hardware and the optimal amount to spend on efficiency grow with the total budget available, although their relative shares may vary. So while an enthusiast may refuse to invest in hardware at all and focus 100% on efficiency, their efforts will still be outclassed by a large company allocating 99% to hardware and 1% to efficiency.
            2. epistasis · · focus · HN ↗
              That&#x27;s simply untrue, they do not have infinite money nor infinite hardware. They have far stricter limits on hardware than they do on development effort to get more out of their finite hardware.
              1. rhdunn · · focus · HN ↗
                They are acting like they have infinite money (huge annual losses [1], [2], [3]) and infinite hardware (aiming to build a large number of data centres with a lot of GPU hardware [1], [3], [4]) even if that isn&#x27;t the case.

                [1] <a href="https:&#x2F;&#x2F;www.morningstar.com&#x2F;stocks&#x2F;anthropics-leaked-financials-reflect-fast-growth-not-2-trillion-valuation" rel="nofollow">https:&#x2F;&#x2F;www.morningstar.com&#x2F;stocks&#x2F;anthropics-leaked-financi...

                [2] <a href="https:&#x2F;&#x2F;www.wheresyoured.at&#x2F;anthropics-profitability-swindle&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.wheresyoured.at&#x2F;anthropics-profitability-swindle...

                [3] <a href="https:&#x2F;&#x2F;fortune.com&#x2F;2026&#x2F;06&#x2F;16&#x2F;openai-financials-leaked-losses-revenue-profit&#x2F;" rel="nofollow">https:&#x2F;&#x2F;fortune.com&#x2F;2026&#x2F;06&#x2F;16&#x2F;openai-financials-leaked-loss...

                [4] <a href="https:&#x2F;&#x2F;www.forbes.com&#x2F;sites&#x2F;paulocarvao&#x2F;2025&#x2F;12&#x2F;06&#x2F;why-openais-ai-data-center-buildout-faces-a-2026-reality-check&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.forbes.com&#x2F;sites&#x2F;paulocarvao&#x2F;2025&#x2F;12&#x2F;06&#x2F;why-open...

        2. holbrad · · focus · HN ↗
          This is such a hilarious comment when we&#x27;ve literally just had Sol 6 and Opus 5.5 drop that are massive wins for efficiency. I mean, how out of touch can you be? Especially if you think power saving and efficiency come from local AI enthusiasts. What a fucking joke.
        3. mrandish · · focus · HN ↗
          &gt; why bother trying to make training and inference efficient when you can just throw hardware&#x2F;money at the problem.

          Sure, for certain bleeding edge research path-finding, any optimization is premature but that&#x27;s tiny compared to deployed end-user product scale. The biggest constraints frontier cloud model providers face are: 1. Getting enough of the hardware they want, 2. Capital, 3. Talent (and #3 is a largely a function of #2 &amp; #1).

          Given the raw economics, it would be foolish to not have efficiency teams fast-following their deployments as closely as possible. We&#x27;re seeing direct and indirect evidence of significant efficiency improvements across all the frontier providers and the rate seems to be increasing. The coming OAI &amp; ANT IPOs hinge on demonstrating they can serve many use cases beyond software dev profitably at what they are willing to pay, which is generally less than software dev. The trend is clearly &quot;the higher volume the use case, the less they&#x27;re willing to pay&quot;.

      2. rpdillon · · focus · HN ↗
        Similarly with my Strix Halo running DSV4. Got a meter for it a couple of weeks back, and, during inference, it&#x27;s pulling about 160W, and I run it about 15 hours a day on a busy day, or about 2.4kW. I&#x27;d say an average day is more like 2 hours for me. It&#x27;s a small impact on my monthly bill. Owning a hot tub is orders of magnitude worse, in my experience.
      3. jmiskovic · · focus · HN ↗
        Inference vs training. It&#x27;s training that takes enormous amount of power, both electrical and the compute. I suspect many many models get simply thrown away because they end up being too low on benchmarks by the time they are done. And some are not released to the public. So we learn only about tiny percentage of trained models and their environmental impact.
        1. mapontosevenths · · focus · HN ↗
          Hey guys, has anyone seen my goalposts?

          More seriously, that needs to be done once and then it can be used by millions of people. Divide the cost by all the users and it&#x27;s trivial. It&#x27;s certainly not enough to lose sleep over.

      4. Tuna-Fish · · focus · HN ↗
        Batching is extremely powerful for optimizing power use. Your card is spending the vast majority of it&#x27;s energy use fetching weights from it&#x27;s memory, and it only gets to use each weight once. Running that same model on a bigger machine with lots of batching gets to amortize the cost of loading a weight across all the parallel requests. N=128 is literally about 50 times more energy-efficient than N=1.
        1. stkdump · · focus · HN ↗
          I don&#x27;t know. It would surprise me. For long agentic work you shouldn&#x27;t need that many streams for the KV cache to become larger than the (active) parameters. At 128 streams the model weights would be neglectable. I am using a 22GB model file, i.e. I have 10GB for context, which isn&#x27;t even the max that the model supports.
    16. christkv · · focus · HN ↗
      Not only an overbuild but one that’s not real as there is a huge difference between talking about building and actually doing the building. Announcements are not the same as execution.
    17. etdznots · · focus · HN ↗
      Inference needs very little compute, it’s mostly IO. hence the low power draw, training is compute-heavy and uses lots of power
      1. apitman · · focus · HN ↗
        Honest question: if that&#x27;s the case why do my GPUs draw their full TDP during inference?
        1. etdznots · · focus · HN ↗
          I am not an expert on why, but my vague answer is you are still running near max clock speeds, and you can significantly downclock your gpu and&#x2F;or set power limits on your GPU without seeing much loss in performance, e.g. you can cap a 3090 to 60% of it’s normal power draw and it will lose like 5% of tg performance and 10% of pp performance (<a href="https:&#x2F;&#x2F;www.reddit.com&#x2F;r&#x2F;LocalLLaMA&#x2F;comments&#x2F;1hg6qrd&#x2F;relative_performance_in_llamacpp_when_adjusting&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.reddit.com&#x2F;r&#x2F;LocalLLaMA&#x2F;comments&#x2F;1hg6qrd&#x2F;relativ...)

          I think theoretically this is still very wasteful with lots of compute that’s getting powered and sitting unused but you are limited either by the firmware or by the GPU architecture from pushing power usage even lower without cratering performance. (no one anticipated the demand for relatively low-compute devices with lots of super fast memory)

      2. Lerc · · focus · HN ↗
        Whil training used a lot of energy, it is quite difficult to comprehend how that measures up in global terms.

        The last numbers I saw for training a frontier model used as much energy as four fully fueled up Boeing Pegasus (of which there are 113)

        The Erin Brokovich data center site describes the use as something like the lifetime use of 5-6 cars, which sounds like a lot when you&#x27;re filling the tank, but hardly anything when you think about how many cars there are.

    18. looofooo0 · · focus · HN ↗
      Jenvons Paradox considered?
    19. gmerc · · focus · HN ↗
      But bro we need to shoot the DC to space because UNLIMITED POWER
    20. simianwords · · focus · HN ↗
      Extremely funny that GPT 6.1 sol is not only cheaper but also more efficient and faster and more performant and therefore better for environment. So they give up on all this just for data sovereignty.
      1. cherioo · · focus · HN ↗
        How do you know sol is more efficient?
        1. simianwords · · focus · HN ↗
          hmm my conviction that labs have high margins
    21. Lerc · · focus · HN ↗
      &gt;With the talk of AI data centers&#x27; impact on the world,

      The global impact of data centers is negligible. There are measurable local impacts. It remains a small percentage of global electricity use. Electricity is less than a quarter of world energy use (increasing that percentage is one of the best things we can do because long term renewables win).

      There are local demands for data centers. Infrastructure strain, from new builds etc.

      A lot of that is dependent on the local situation, for instance water use in any area that performs regular irrigation will have additional use of water, but the relative quantities lean heavily towards irrigation.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.