‹ BackHN Continuity

Thread

How GLM built its own inference infrastructure

411 points · 285 comments · whiteros_e

  1. embedding-shape · · focus · HN ↗
    I was gonna ask how people found their coding plans, and realized, have they massively ramped up the prices? Seems the middle plan is ~$80/month now, didn't that used to be like $20/month? Cheapest plan is ~$20/month currently.

    They must have hit really hard scaling limits if the prices were hiked so much so quickly.

    1. Daviey · · focus · HN ↗
      I paid $360 annual for Max plan and currently averaging about 1BN tokens a day with their frontier GLM-5.3 model. This was clearly unsustainable for them and they've dropped this package.
      1. world2vec · · focus · HN ↗
        1 billion tokens a day?!! I've done a lot of work these past 2 weeks with GLM-5.3. Like, a lot. And I've just passed 300 million tokens in total.

        Can I ask where are you using all those tokens?

        1. _0ffh · · focus · HN ↗
          Well, there's essentially two major ways to use these models: Pair programming or fully autonomous fire-and-forget code generation. The second strategy needs essentially zero input, so the number of tokens you can blow is practically only limited by API speed.
          1. rubslopes · · focus · HN ↗
            There's also a third way that can spend the most tokens: if the AI is used as part of the product, and not just a tool to build the product.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.