‹ BackHN Continuity

Thread

How GLM built its own inference infrastructure

411 points · 285 comments · whiteros_e

  1. zicohacks · · focus · HN ↗
    US chip export restrictions may actually be an advantage for China's AI Infrastructure. Chinese companies are forced to speed up developing their own AI chips
    1. menaerus · · focus · HN ↗
      It was evident that this will happen.

      > Compared with our initial baseline on the same hardware, we achieved a 3× improvement in end-to-end serving performance, reaching hardware efficiency and per-token cost comparable to mainstream NVIDIA GPUs. This demonstrates that Chinese chips can support frontier-model inference efficiently and economically at scale.

      1. verdverm · · focus · HN ↗
        you can also derive some stats from the ~10T tokens a day on 100k devices, 100M / device / day, but then one has to account for the multi-gpu model size, and I need coffee before I go there
        1. a34729t · · focus · HN ↗
          I mean all the database stuff is obvious low hanging fruit for inference engines.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.