‹ BackHN Continuity

Thread

How GLM built its own inference infrastructure

411 points · 285 comments · whiteros_e

  1. dada216 · · focus · HN ↗
    We built a complete production-grade inference service from scratch on a cluster of more than 100,000 Chinese-made AI accelerators. All production inference for GLM-5.3-Flash runs on this system.
    1. freakynit · · focus · HN ↗
      Most of the people had kinda guessed this when they decided to provide 100 trillion tokens for free.
      1. gpugreg · · focus · HN ↗
        It wasn&#x27;t a secret either. They blogged about it last month: <a href="https:&#x2F;&#x2F;z.ai&#x2F;blog&#x2F;glm-5.3-flash#:~:text=Serving%20at%20Scale%20on%20Chinese%20AI%20Chips" rel="nofollow">https:&#x2F;&#x2F;z.ai&#x2F;blog&#x2F;glm-5.3-flash#:~:text=Serving%20at%20Scale...
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.