One month coding with GLM 5.3 Flash
Thread
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
One month coding with GLM 5.3 Flash
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
ThibWeb · · focus · HN ↗
cvburgess · · focus · HN ↗
gunalx · · focus · HN ↗
ThibWeb · · focus · HN ↗
RussianCow · · focus · HN ↗
sampullman · · focus · HN ↗
lnenad · · focus · HN ↗
samtheprogram · · focus · HN ↗
ThibWeb · · focus · HN ↗
nowittyusername · · focus · HN ↗
RussianCow · · focus · HN ↗
sampullman · · focus · HN ↗
joshheitzman · · focus · HN ↗
throw930rmdkdk · · focus · HN ↗
eikenberry · · focus · HN ↗
esafak · · focus · HN ↗
verdverm · · focus · HN ↗
Geof25 · · focus · HN ↗
verdverm · · focus · HN ↗
It's my favorite model family to interact with, it's prose is the best imo, it makes me laugh from time-to-time (like when it said it would "crib" some code from another project, lul)
hunter2026 · · focus · HN ↗
[dead]
esafak · · focus · HN ↗
gunalx · · focus · HN ↗
jminnl · · focus · HN ↗
What kind of hardware and what particular quant?
aktenlage · · focus · HN ↗
> Nonetheless, it’s a good reminder to be careful with model selection and with agentic patterns. We could have achieved similar results for most likely 5x less cost with not that much more effort. Lessons learned! We need to budget for this, and be more careful. Could have seen it coming, but now we know.
I don't get it. Why was it wrong? Which one would have been better? What was the lesson and how could you have foreseen it?
ThibWeb · · focus · HN ↗
gpugreg · · focus · HN ↗
ThibWeb · · focus · HN ↗
epistasis · · focus · HN ↗
> That model’s usage was well within our budget ($68, about 4kWh of energy use / 365 grams of carbon emissions).
The energy cost is literally 1% of the total cost. For context, 4kWh of energy would drive you about 15 miles in an EV, about half of the average person's daily driving miles. It's boiling 10 gallons of water.
With the talk of AI data centers' impact on the world, you'd think this would be 10x to 100x the amount of energy in order to get the effects they're using here.
My takeaway: the AI data center buildout is an overbuild probably at least as large as the fiber buildout that left us with so much dark fiber. If not even bigger. The only thing that will save the economy is the inability of NVIDIA and chip fabs to produce enough chips to match the buildout planned.
ThibWeb · · focus · HN ↗
CuriouslyC · · focus · HN ↗
pphysch · · focus · HN ↗
CuriouslyC · · focus · HN ↗
lukan · · focus · HN ↗
So how many companies/individuals pay a random spam bot contacting them 100€ to upgrade their website? I guess start with targeting gullible and trying to scam them will be the more successful experiment for them. (If run unrestricted)
mannanj · · focus · HN ↗
MisterMunchkin · · focus · HN ↗
CuriouslyC · · focus · HN ↗
monkpit · · focus · HN ↗
Then why would any AI company exist when they could just use their money to buy tokens from another AI company and make money for zero effort? There would be no incentive to be a provider.
Not to mention inflation would grow to match or outpace the rate you could earn on these guaranteed AI gains.
rhyperior · · focus · HN ↗
cyanydeez · · focus · HN ↗
gunalx · · focus · HN ↗
epistasis · · focus · HN ↗
In the equilibrium, profits fall to zero in capitalism, but there's never equilibrium, everything is changing. Profit comes from insights that others haven't seen, from access to opportunities that others don't have, or through monopolistic control generating economic rents.
For that AI to make money, it has to have a harness that gives it one of the above things, which does seem possible.
swiftcoder · · focus · HN ↗
But strictly time-limited, because as soon as someone else figures out what you are doing, there's no moat to prevent them from telling their AI to do it too.
Markets reach equilibrium very fast when everyone has access to the exact same tools and information.
a3dds · · focus · HN ↗
Take a vacation.
rapatel0 · · focus · HN ↗
skew-aberration · · focus · HN ↗
rapatel0 · · focus · HN ↗
It's a finger in the wind approximation with lots of confounders, but historically it seems to work out that way for most circuit things.
233mhz · · focus · HN ↗
My car engine is a miracle of efficiency compared to the stuff we had a hundred years ago, but overall cars pollute more than back then
driverdan · · focus · HN ↗
spiffytech · · focus · HN ↗
driverdan · · focus · HN ↗
ThibWeb · · focus · HN ↗
oblio · · focus · HN ↗
So potentially increasing energy usage even more, especially with likely token usage acceleration, is hardly reason for celebration.
driverdan · · focus · HN ↗
nicoburns · · focus · HN ↗
I don't
> Where I live isn't relevant,
But where you live does matter. Average monthly household energy usage in the EU is 200-400 kWh. In the US it's more like 800-900kWh. So what you consider "normal energy usage" may vary considerably by where you live.
At 200-400kWh total usage, 30kwh starts to look a lot more significant. So you should at least consider the possibility that it's not that energy usage from AI is insignificant, but that your existing energy usage is wasteful.
oblio · · focus · HN ↗
My point isn't to attack you personally (I'd rather attack your country's disastrous energy policies) but to point out that LLMs will constitute a huge increase in our energy consumption at a time where we haven't transitioned to renewable energy and we're probably 20-30 away from it.
childintime · · focus · HN ↗
233mhz · · focus · HN ↗
And the average European house uses like 3000kwh.... per year. We're already wasting insane amount of energy, adding 2-10% more is definitely huge
killingtime74 · · focus · HN ↗
ThibWeb · · focus · HN ↗
epistasis · · focus · HN ↗
My entire point is that 99% of the dollar cost of running these models goes to things other than the GPU power. The capex cost to building cost to GPU cost to storage/networking/chasses/wiring plus the other operation costs dwarf the electricity.
That's massive and the constraints on fabs, etc. will drive the amount of the AI build far more than energy availability.
smartmic · · focus · HN ↗
I have the feeling China is somehow ahead when it comes to energy (and cost) efficiency for AI usage.
eli · · focus · HN ↗
jjcm · · focus · HN ↗
a.) we’re supply constrained
b.) only 3% of households pay for AI
Inference amounts will continue to grow heavily.
timmmmmmay · · focus · HN ↗
DragonStrength · · focus · HN ↗
api · · focus · HN ↗
epistasis · · focus · HN ↗
Add in 1) all job destruction that Sam Altman and other leaders tout, 2) Musk operating an environmental disaster in the most offensive way possible for xAI, and 3) general fear bait a new technology that has such uncertain consequences, and I'm surprised the backlash isn't stronger.
Go talk to community members about any change in the buildings in their area and the data center backlash fits right in line with normal responses.
duskdozer · · focus · HN ↗
cyanydeez · · focus · HN ↗
I can't believe the TAM requires more than it at a 2x speed up. OpenAI, Anthropic built models whose only purpose now is things like research and defense.
And by defense, obviously, in America, we mean war. killing, etc.
ericd · · focus · HN ↗
cyberrock · · focus · HN ↗
Though I do suspect that golf course owners in Arizona and alfalfa farmers in California must be ecstatic.
zozbot234 · · focus · HN ↗
epistasis · · focus · HN ↗
As we electrify and decarbonize, we are looking at going to 1000-1500GW average load (clean energy sources are 2x-6x more energy efficient than fossil fuels at delivering energy services , so looking at the primary energy flows here you'll not only see elimination of the "rejected energy" category, but heating by fossil fuels gets counted as ~100% effect when heat pumps are 200%-600% effect on those terms <a href="https://flowcharts.llnl.gov/" rel="nofollow">https://flowcharts.llnl.gov/ )
Going to 50GW of AI would be an amazing jump for which we don't have fab capacity anytime.
gchamonlive · · focus · HN ↗
hedora · · focus · HN ↗
So, assume 8 of those, and it’s a space heater per house. I have never been disturbed by my neighbor’s space heater.
The problem is centralization, not the absolute energy usage. (And also that LLMs are trending to 100x more efficient than the data center sizing assumed).
gchamonlive · · focus · HN ↗
epistasis · · focus · HN ↗
Centralization is a huge huge environmental win, far far more than even cloud computing was compared to tons of inefficient, under utilized racks spread though our office buildings.
z0r · · focus · HN ↗
rpdillon · · focus · HN ↗
hypfer · · focus · HN ↗
In a vacuum, yes. Assuming everyone is a good actor.
rpdillon · · focus · HN ↗
epistasis · · focus · HN ↗
Calculate the scale. Add it up. Do it.
Look at the numbers and come back to me.
ls612 · · focus · HN ↗
gchamonlive · · focus · HN ↗
gchamonlive · · focus · HN ↗
Fidelix · · focus · HN ↗
gchamonlive · · focus · HN ↗
simianwords · · focus · HN ↗
gchamonlive · · focus · HN ↗
233mhz · · focus · HN ↗
gchamonlive · · focus · HN ↗
epistasis · · focus · HN ↗
If somebody is afraid of speaking their mind, because the are afraid of being labelled an opponent to progress, that's some personal issue to work through.
It is popular and encouraged to be skeptical of AI, there no social opprobrium about it, unless you're in an extremist political cult, in which case you probably call it SI.
gchamonlive · · focus · HN ↗
redanddead · · focus · HN ↗
I never understand the mindset that leads to these massive leaps in reasoning. I was recruiting a guy, he was a VC GP, he says the same thing. I think he’s wrong. The more we use the more we want to use.
Do we seriously think our species will jinx the Kardashev energy requirements
epistasis · · focus · HN ↗
Every massive technological leap that have big private buildouts results in massive overbuild. It's the nature of FOMO when a big civilization changing tech gets introduced.
Will AI change everything? Are too many data centers planned? The answer is almost certainly yes to both.
redanddead · · focus · HN ↗
rpdillon · · focus · HN ↗
swiftcoder · · focus · HN ↗
Yes, but the lag there is decade-long, and a bunch of heavily fibre-invested companies went bust before anyone saw the upside.
In the same vein, it's very possible that early movers-and-shakers in the LLM space will overextend and end up bankrupt before LLMs find a long-term profitable niche
rpdillon · · focus · HN ↗
scottcha · · focus · HN ↗
Regarding total costs relative to the pure energy costs it is multiple orders of magnitude different but also realize in the datacenter the energy is the pure commodity while almost every other component has huge margins driven by lack of supply. I do think over time this might get closer together (more competition on HW might lower margins) while energy might become more of a bottle neck (raising the energy prices).
bix6 · · focus · HN ↗
Big corps bullied (leveraged) communities and hid it and lied or withheld info. They refuse to share basic numbers on planned builds. I thought they had PR teams.
polytely · · focus · HN ↗
scottcha · · focus · HN ↗
lhl · · focus · HN ↗
These differ by model and set of tasks, but for GLM-5.3 Flash, the best $/pass was medium effort. Max did get about 5% higher scores, but at >2X the $/pass.
ThibWeb · · focus · HN ↗
scottcha · · focus · HN ↗
teaearlgraycold · · focus · HN ↗
Well there's also training to consider. But yeah, most of what I see online about datacenters is nonsense. People worried about water use as a top concern are misinformed.
Look, I absolutely get it if you're worried your boss is just waiting for the day to replace you with an LLM. Fucking get organized with your fellow laborers instead of getting distracted by datacenter water use. Call your politicians to rein in billionaires. If anything the ownership class is probably directing online discourse towards electricity and water user to keep people away from class-focused debates.
etdznots · · focus · HN ↗
teaearlgraycold · · focus · HN ↗
stkdump · · focus · HN ↗
Having said that, seeing the incredible progress of models throughout this year, I also strongly believe that the planned buildout is overeager. Even I with my gaming hardware often run out of instructions to give. And the smarter the models get that I can run, the less I will be able to saturate my hardware. Is it because of my lack of creativity of which kinds of tasks I can give to AI? Maybe a bit, but currently I can't believe that I am that far away.
rhdunn · · focus · HN ↗
I think the power savings and efficiencies (that the large companies have also benefited from) have come from 2 areas:
1. open source and local AI enthusiasts -- think things like llama.cpp, quantization, etc.
2. Chinese labs and other smaller/research companies like Mistral that are using constrained hardware -- see the various advancements in the various models to reduce compute complexity such as mixture of experts [1], sharing key/value data between a group of layers, etc.
[1] Though the original idea for mixture of experts comes from a 1991 research paper (<a href="https://huggingface.co/blog/moe" rel="nofollow">https://huggingface.co/blog/moe), so maybe a third area is research from Universities, etc.
yorwba · · focus · HN ↗
darkwater · · focus · HN ↗
yorwba · · focus · HN ↗
epistasis · · focus · HN ↗
rhdunn · · focus · HN ↗
[1] <a href="https://www.morningstar.com/stocks/anthropics-leaked-financials-reflect-fast-growth-not-2-trillion-valuation" rel="nofollow">https://www.morningstar.com/stocks/anthropics-leaked-financi...
[2] <a href="https://www.wheresyoured.at/anthropics-profitability-swindle/" rel="nofollow">https://www.wheresyoured.at/anthropics-profitability-swindle...
[3] <a href="https://fortune.com/2026/06/16/openai-financials-leaked-losses-revenue-profit/" rel="nofollow">https://fortune.com/2026/06/16/openai-financials-leaked-loss...
[4] <a href="https://www.forbes.com/sites/paulocarvao/2025/12/06/why-openais-ai-data-center-buildout-faces-a-2026-reality-check/" rel="nofollow">https://www.forbes.com/sites/paulocarvao/2025/12/06/why-open...
holbrad · · focus · HN ↗
mrandish · · focus · HN ↗
Sure, for certain bleeding edge research path-finding, any optimization is premature but that's tiny compared to deployed end-user product scale. The biggest constraints frontier cloud model providers face are: 1. Getting enough of the hardware they want, 2. Capital, 3. Talent (and #3 is a largely a function of #2 & #1).
Given the raw economics, it would be foolish to not have efficiency teams fast-following their deployments as closely as possible. We're seeing direct and indirect evidence of significant efficiency improvements across all the frontier providers and the rate seems to be increasing. The coming OAI & ANT IPOs hinge on demonstrating they can serve many use cases beyond software dev profitably at what they are willing to pay, which is generally less than software dev. The trend is clearly "the higher volume the use case, the less they're willing to pay".
rpdillon · · focus · HN ↗
jmiskovic · · focus · HN ↗
mapontosevenths · · focus · HN ↗
More seriously, that needs to be done once and then it can be used by millions of people. Divide the cost by all the users and it's trivial. It's certainly not enough to lose sleep over.
Tuna-Fish · · focus · HN ↗
stkdump · · focus · HN ↗
christkv · · focus · HN ↗
etdznots · · focus · HN ↗
apitman · · focus · HN ↗
etdznots · · focus · HN ↗
I think theoretically this is still very wasteful with lots of compute that’s getting powered and sitting unused but I think at that point you are limited either by the firmware or by the GPU architecture from pushing power usage even lower (no one anticipated the demand for relatively low-compute devices with lots of super fast memory)
Lerc · · focus · HN ↗
The last numbers I saw for training a frontier model used as much energy as four fully fueled up Boeing Pegasus (of which there are 113)
The Erin Brokovich data center site describes the use as something like the lifetime use of 5-6 cars, which sounds like a lot when you're filling the tank, but hardly anything when you think about how many cats there are.
looofooo0 · · focus · HN ↗
gmerc · · focus · HN ↗
simianwords · · focus · HN ↗
cherioo · · focus · HN ↗
simianwords · · focus · HN ↗
Lerc · · focus · HN ↗
The global impact of data centers is negligible. There are measurable local impacts. It remains a small percentage of global electricity use. Electricity is less than a quarter of world energy use (increasing that percentage is one of the best things we can do because long term renewables win).
There are local demands for data centers. Infrastructure strain, from new builds etc.
A lot of that is dependent on the local situation, for instance water use in any area that performs regular irrigation will have additional use of water, but the relative quantities lean heavily towards irrigation.
disiplus · · focus · HN ↗
raggi · · focus · HN ↗
The second reason appeared to be simply "because we chose not to". The post seems to be pretty much content-less in any practical sense. I clicked on it because I do quite like this models average performance and I was hoping to see some kind of review content.
ThibWeb · · focus · HN ↗
desmondl · · focus · HN ↗
He said his experiment was a failure because:
1. He accidentally spent 450M tokens vibe coding with the wrong model, instead of GLM 5.3 Flash.
2. When he used GLM 5.3 Flash, it was sometimes slow. So he switched to other models (Deepseek / Qwen) instead. His guess to why it was slow: GLM 5.3 Flash was so good that the providers were congested.
3. He still needed to use other models besides GLM 5.3 Flash, for R&D and benchmarking.
His takeaways from doing the experiment were:
1. Measure local usage more.
2. Experiment with agent orchestration, with bounded goals.
3. Don't count other models that are used for R&D.
4. Play with Jev.
5. Include experiments with flagship models to compare with cheap open models.
His conclusion about GLM 5.3 Flash: Probably viable for day to day work, but he'll have more thoughts next month.
jensb1 · · focus · HN ↗
ydj · · focus · HN ↗
monksy · · focus · HN ↗
surgical_fire · · focus · HN ↗
It's an excellent workhorse. When I am running out of my GLM quota I switch GLM-5.3-flash to DS-4.1-flash.
finnjohnsen2 · · focus · HN ↗
surgical_fire · · focus · HN ↗
finnjohnsen2 · · focus · HN ↗
surgical_fire · · focus · HN ↗
I find that it makes coding and review at the same time less prone to errors and cheaper
drob518 · · focus · HN ↗
monksy · · focus · HN ↗
drob518 · · focus · HN ↗
girvo · · focus · HN ↗
vardalab · · focus · HN ↗
girvo · · focus · HN ↗
The legacy plan I have is so good as to be basically unlimited usage for my workloads, so I’m kind of stuck with it til they stop renewing it haha
drob518 · · focus · HN ↗
pmoriarty · · focus · HN ↗
[1] - <a href="https://www.youtube.com/watch?v=P4dTq4X8bqk" rel="nofollow">https://www.youtube.com/watch?v=P4dTq4X8bqk
benjiro29 · · focus · HN ↗
1. Models like DeepSeek V4.1 Flash are much cheaper on DeepSeek their API directly because of the cache handeling is better. Neuralwatt can hit up to 98% but DeepSeek can do 99.x... That may not sound like a big difference but it quickly widens the gap on long tasks to grow 2x a 3x in price. DeepSeek their cache handeling is S-tier (with a ton of features, for instance 24h caching).
2. The same issue is also present if you compare GLM 5.3 Flash with z.ai vs Neuralwatt. Its just way more cheaper from the source, then from Neuralwatt.
3. The energy numbers from Neuralwatt are ... to be taken with a ton of salt. Past energy numbers had the same models (for instance) GLM 5.2 up to 6x cheaper in energy usage, then after they "fixed" issues with the energy numbers. In reality, those energy numbers are just a different form of billing, but not a actual representation of the energy usage of AI models. Things like profits are inside those energy numbers. So seeing 4kWH used for a model, does not mean that it uses 4Kwh.
eli · · focus · HN ↗
I think NW's profits are mostly between what they pay for electricity and what they charge you for electricity.
ThibWeb · · focus · HN ↗
sheepscreek · · focus · HN ↗
Makes me appreciate my ChatGPT subscription. I’ve had multiple days between 1B-2B tokens (now less so, models have indeed become token efficient) and regularly in the > 100M range. Even then, $150 sounds excessive. I wonder if their cache is getting nuked for some reason, or maybe they decide to use Cerebras that doesn’t subsidize cached tokens.
ThibWeb · · focus · HN ↗
RussianCow · · focus · HN ↗
sheepscreek · · focus · HN ↗
<a href="https://imgur.com/a/mSXGzzJ" rel="nofollow">https://imgur.com/a/mSXGzzJ
fallinditch · · focus · HN ↗
GLM 5.3 flash has been good as my default profile Hermes bot, after I readjusted its memory to point to a couple of key skills.
I got some great coding results with GLM 5.2, and 5.3 Flash is supposedly almost as good, so I will be trying it out soon for day to day tasks as the post advises.
UncleOxidant · · focus · HN ↗
buildbot · · focus · HN ↗
UncleOxidant · · focus · HN ↗
hunter2026 · · focus · HN ↗
[dead]
consumer451 · · focus · HN ↗
I am a huge Anthropic fan, USA fan, but are we cooked with AI? Sorry SI, that's the important thing.
hgoel · · focus · HN ↗
The direct excitement regarding open models is that many of these Flash Mixture-of-Expert models run reasonably well on hardware a tech employee in the West, and businesses in less affluent countries, can realistically afford.
The indirect excitement is that the models are so efficient, so cloud prices also end up being very low.
You don't see the same scale of excitement surrounding the open weight trillion+ parameter models because, while it's neat they're open weight, it doesn't mean a lot if you need $50k worth of computers to just barely run them.
consumer451 · · focus · HN ↗
However, we are a deeply stupid species. Based on multiple previous conversations on this website, it appears that for example, z.ai hosting is not a big deal.
Yes, anyone doing real due diligence on client data will face reality, maybe. However, given the long tail of external consultancies, the CCP is going to eat it all due to our outsourced laziness. Shareholder value, amirite!!?
This is how we lost our manufacturing base. Why would the token manufacturing base be any different? As far as I can tell, we have gotten even stupider in the last few years. I guess changing the name to SI and tariffs on Canada and the EU will solve these problems.
consumer451 · · focus · HN ↗
hgoel · · focus · HN ↗
Personally I am of the opinion that the Western frontier companies are increasingly evil in a "Those who torment us for our own good will torment us without end, for they do so with the approval of their own conscience." sense.
But, as far as the diversity of frontier research companies in the West is concerned (I believe you overlooked Google and Facebook), I think the issue is that they are all very flush with cash, don't care much for competing fairly, and are deeply embedded with the government. So, competitors mostly die in the cradle, any that make it can be purchased, and any that refused can be crushed by the government.
zozbot234 · · focus · HN ↗
This is a rather outdated POV. With increasingly pervasive use of SSD offload, there nothing particular stopping you from running even the largest open models on ordinary local hardware. Sure, it will be really slow, but if you need the smarts for e.g. a one-off planning role it's a no-brainer.
consumer451 · · focus · HN ↗
The price one pays for this is that lower than frontier models are dumber. At the current pace of advancement, is that a good idea? Isn't it better to just go with with ZDR contracts, and not lose in the software product game, which now moves at relativistic speed?
hgoel · · focus · HN ↗
For now though, I would question if the difference in smarts is large enough to justify tolerating a generation speed measured in seconds per token compared to spending more turns refining a plan with a flash model.
Tepix · · focus · HN ↗
hgoel · · focus · HN ↗
Though, Qwen3.8-Flash-Next is very close to that level while requiring fewer resources to run, so I'm really looking forward to Qwen4.
vardalab · · focus · HN ↗
nxobject · · focus · HN ↗
everlastingemai · · focus · HN ↗
lin7c · · focus · HN ↗
[dead]
theunclej · · focus · HN ↗
[dead]
andai · · focus · HN ↗
My two favorite graphs are AA's A Index vs Time Per Task and AA Index vs Cost Per Task.
<a href="https://artificialanalysis.ai/#intelligence-comparison-tabs" rel="nofollow">https://artificialanalysis.ai/#intelligence-comparison-tabs
<a href="https://artificialanalysis.ai/?intelligence-comparison=intelligence-vs-time-per-task#intelligence-comparison-tabs" rel="nofollow">https://artificialanalysis.ai/?intelligence-comparison=intel...
ThibWeb · · focus · HN ↗
ericpauley · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
cloudengineer94 · · focus · HN ↗
I created it's his own bot profile even...
[deleted] · · focus · HN ↗
[deleted]
shieldagent · · focus · HN ↗
Aboutplants · · focus · HN ↗
dash-44 · · focus · HN ↗
There was 4.3 months between Opus 4.7 and GLM 5.3 flash. Shorten your timelines ;).
dash-44 · · focus · HN ↗
This is just such a bad opening statement, almost pure clickbait. It makes it sound like the rest of the article is going to be on how it was a mistake to go with cheap models because their performance isn't good enough etc; the same usual stuff we see when someone tries to replace a frontier model with a cheap/flash model.
But the reality is that GLM 5.3 flash if a fantastic model; very cheap/performant and really a new era for this type of model, and the blog is just about some completely unrelated mistakes in development processes. Give us our time back.
kelnos · · focus · HN ↗