‹ BackHN Continuity

Thread

Turning GLM-5.3-Flash into a Jev-like decision model

138 points · 59 comments · flxflx

  1. m4y0u · · focus · HN ↗
    My question is why not use Jev instead? It's faster and cheaper.
    1. kylecazar · · focus · HN ↗
      There's some speculation that Jev is an open weight model with novel post-training (RLCD). So, if these folks have competitive accuracy with just the base model, it may raise some questions about the necessity of Jev's architecture. You generally don't want to find yourself competing only on price.

      Fyi, I haven't tested this yet.

      1. janalsncm · · focus · HN ↗
        > it may raise some questions about the necessity of Jev's architecture

        When I hear “architecture” I am thinking number of parameters and latency.

        When I hear “accuracy” I think training recipe, data, and (later) number of parameters.

        So when you say that Jev’s architecture may not be necessary, the evidence I expect to see is comparable quality at comparable latency. Not equal quality at 2x latency and 4x the cost.

        1. kylecazar · · focus · HN ↗
          The post claims comparable quality at comparable latency (the difference in the latter only stemming from serving region).

          Hence my note about the risk of cost being the only moat.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.