‹ BackHN Continuity

Thread

How GLM built its own inference infrastructure

411 points · 285 comments · whiteros_e

  1. Argonautlabs · · focus · HN ↗
    Different angle on the same model: the full GLM-5.3 (744B MoE, 4-bit experts, 434 GB on disk) runs on a single MacBook Pro M5 Max with 128 GB by streaming the experts from NVMe SSDs instead of keeping them in memory.

    One drive gives about 2 tok/s; striped across four drives it reaches 3.5 tok/s with byte-identical output, and our best internal build with a not-yet-published patch does 4.2.

    Method and numbers: <a href="https:&#x2F;&#x2F;github.com&#x2F;argonautlabsai&#x2F;argodrive" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;argonautlabsai&#x2F;argodrive (built on antirez&#x2F;ds4).

    1. tipsytoad · · focus · HN ↗
      seems unusably slow, and is this for short context?
      1. Argonautlabs · · focus · HN ↗

        [dead]

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.