‹ BackHN Continuity

Thread

How GLM built its own inference infrastructure

411 points · 285 comments · whiteros_e

  1. Havoc · · focus · HN ↗
    Interesting that the tone of announcements between US and Chinese providers is converging.

    GLM has in the past been more technical rather than speculation about future development on RSI etc.

    Also curious whether those 100k accelerators are entirely locally made. If that's genuinely end to end on all components including lithography, memory, design etc then that is quite a feat.

    1. dude250711 · · focus · HN ↗
      Any details on the latest approach to distillation would also be very interesting.
      1. Schlagbohrer · · focus · HN ↗
        I am surprised at the lack of open-weights models in the >35B, but <200B range. I keep thinking about devices like the NVIDIA Spark and AMD Ryzen Halo, which have their 128GB of combined memory, but there are so few models made for that range. Nearly all the open weights distillations are for larger customer bases with <24GB VRAM.
        1. bitexploder · · focus · HN ↗
          Qwen Flash Next 3.8 … even at 3 bit quant it is very solid.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.