‹ BackHN Continuity

Thread

How GLM built its own inference infrastructure

411 points · 285 comments · whiteros_e

  1. embedding-shape · · focus · HN ↗
    I was gonna ask how people found their coding plans, and realized, have they massively ramped up the prices? Seems the middle plan is ~$80/month now, didn't that used to be like $20/month? Cheapest plan is ~$20/month currently.

    They must have hit really hard scaling limits if the prices were hiked so much so quickly.

    1. bbor · · focus · HN ↗
      It's hard to know, since no one advertises the actual token limits (partially cause they're prolly complex / adaptive). So it seems much more likely that they just offer different pricing tiers than you're used to. Like, the $80 plan is still ~$80 of subscription quota, regardless of what else is offered.

      For [API usage](<a href="https:&#x2F;&#x2F;openrouter.ai&#x2F;z-ai&#x2F;glm-5.3-flash#providers" rel="nofollow">https:&#x2F;&#x2F;openrouter.ai&#x2F;z-ai&#x2F;glm-5.3-flash#providers) they charge a bit more than the very cheapest providers of GLM-5.3-Flash, but not so much that a big price difference would make sense.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.