‹ BackHN Continuity

Thread

How GLM built its own inference infrastructure

411 points · 285 comments · whiteros_e

  1. a11r · · focus · HN ↗
    Pretty impressive to see the amount of performance they can squeeze out of the same hardware. I suspect the same process will play out for all combinations of LLMs, inference providers and hardware stacks. This should bring down the cost of inference for the providers by an order of magnitude in the next year and lead to fantastic margins for inference providers.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.