‹ BackHN Continuity

Thread

Turning GLM-5.3-Flash into a Jev-like decision model

138 points · 59 comments · flxflx

  1. ricardobeat · · focus · HN ↗
    Everyone is doing this to emulate Jev, but...

    I took a random book excerpt with 23,000 words (±30k input tokens) and used it as context. Jev still responds in 800ms, sometimes 500ms. That's in the neighbourhood of 20-50,000 tok/s prefill, which is obviously not possible with normal LLMs, not even Cerebras is this fast.

    1. manojlds · · focus · HN ↗
      Do we have a reliable way to count tokens for Jev yet btw?
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.