‹ BackHN Continuity

Thread

How GLM built its own inference infrastructure

411 points · 285 comments · whiteros_e

  1. embedding-shape · · focus · HN ↗
    I was gonna ask how people found their coding plans, and realized, have they massively ramped up the prices? Seems the middle plan is ~$80/month now, didn't that used to be like $20/month? Cheapest plan is ~$20/month currently.

    They must have hit really hard scaling limits if the prices were hiked so much so quickly.

    1. Daviey · · focus · HN ↗
      I paid $360 annual for Max plan and currently averaging about 1BN tokens a day with their frontier GLM-5.3 model. This was clearly unsustainable for them and they've dropped this package.
      1. world2vec · · focus · HN ↗
        1 billion tokens a day?!! I've done a lot of work these past 2 weeks with GLM-5.3. Like, a lot. And I've just passed 300 million tokens in total.

        Can I ask where are you using all those tokens?

        1. tokai · · focus · HN ↗
          300M for two weeks is surprisingly low. What are you doing that need so few tokens?
          1. world2vec · · focus · HN ↗
            It's not my main model (that would be Fable 5.1 Extra) but it's been doing agent-driven search and optimisation of a cross-trading ranking model (it's for work).
            1. disiplus · · focus · HN ↗
              I would suggest you to hook fable or 5.6 to check it regularly and its work because it gets lost easily on stuff it was not trained on. I'm doing some custom inference engine optimization and it's a workhorse but it can easily lose its way and if you don't recheck it you will get wrong answers in the end.
              1. world2vec · · focus · HN ↗
                Yeah that's what I already do. Fable writes the plan and checks things at certain milestones. Otherwise it does get lost indeed.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.