‹ BackHN Continuity

Thread

How GLM built its own inference infrastructure

411 points · 285 comments · whiteros_e

  1. 9cb14c1ec0 · · focus · HN ↗
    Given the huge amount of money being spent on AI chips in the US, what prevents US AI labs from doing the same level of software optimization? It could be a solve for some of the capacity constraints.
    1. kingstnap · · focus · HN ↗
      They do, when Luna got 5x cheaper it was directly attributed to some unknown % inference optimization.

      US labs are quite cut throat about dealing with stuff costing them money (inference). This sort of engineering excellence doesn't always feel that way because they are simultaneously quite lax about stuff costing other people money.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.