‹ BackHN Continuity

Thread

Turning GLM-5.3-Flash into a Jev-like decision model

138 points · 59 comments · flxflx

  1. ricardobeat · · focus · HN ↗
    Everyone is doing this to emulate Jev, but...

    I took a random book excerpt with 23,000 words (±30k input tokens) and used it as context. Jev still responds in 800ms, sometimes 500ms. That's in the neighbourhood of 20-50,000 tok/s prefill, which is obviously not possible with normal LLMs, not even Cerebras is this fast.

    1. mmastrac · · focus · HN ↗
      That's not true. I ran Cerebras as an experimental ultrafast Jev and it was faster.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.