We built a complete production-grade inference service from scratch on a cluster of more than 100,000 Chinese-made AI accelerators. All production inference for GLM-5.3-Flash runs on this system.
It wasn't a secret either. They blogged about it last month: <a href="https://z.ai/blog/glm-5.3-flash#:~:text=Serving%20at%20Scale%20on%20Chinese%20AI%20Chips" rel="nofollow">https://z.ai/blog/glm-5.3-flash#:~:text=Serving%20at%20Scale...
dada216 · · focus · HN ↗
freakynit · · focus · HN ↗
gpugreg · · focus · HN ↗