‹ BackHN Continuity

Thread

Turning GLM-5.3-Flash into a Jev-like decision model

138 points · 59 comments · flxflx

  1. m4y0u · · focus · HN ↗
    My question is why not use Jev instead? It's faster and cheaper.
    1. andrewchambers · · focus · HN ↗
      These questions are answered by the OP (Same speed, image support) - additionally, GLM is open weight.
    2. kylecazar · · focus · HN ↗
      There's some speculation that Jev is an open weight model with novel post-training (RLCD). So, if these folks have competitive accuracy with just the base model, it may raise some questions about the necessity of Jev's architecture. You generally don't want to find yourself competing only on price.

      Fyi, I haven't tested this yet.

      1. dcss_gardener · · focus · HN ↗
        I mean just from what's known of the funding and timeline it pretty much has to be based on open weights.

        But it is likely more than just a fine tune + novel training. At the very least the LM head is swapped out for a classifier one and then or also idk, bidirectional attention for the encoding pass I'm out of my depth at this point and will stop guessing. The training is probably where they have the biggest moat though, not that it's necessarily huge.

        I have a project that fits jev as advertised almost comically well and I've been playing with it, and the various hacks and open versions. Jev doesn't necessarily perform better overall but it is quite different. It's sensitive to prompt phrasing in ways the others aren't, it's easy to generate questions where all the other models cluster in confidence but jev is an outlier. Not necessarily more correct, but it does feel like it's getting its answers in a different way.

        I'm guessing just as much as anyone else but I've been spending a ton of time on this the last couple weeks, it landed right when I was most ready to dig into it.

        1. ComputerGuru · · focus · HN ↗
          Your last bit about it being different is a known issue with models that have been trained on purely synthetic data, no?
          1. dcss_gardener · · focus · HN ↗
            Sure but I still don’t see any reason to think it’s not based on an open weight model. They don’t even make that claim as far as I know.
            1. ComputerGuru · · focus · HN ↗
              Yeah, I’m similarly not sure how to square away claims of “trained 100% exclusively on synthetic data” with the fairly obvious fact that it’s derived from an open model, almost all of which are trained on a mix of scraped and data and distilled traces.
      2. Fordec · · focus · HN ↗
        Also, while it's clearly got a lot of training on some use cases, others that probably weren't in the training set have worse good decision rates than a random number generator. If you can rebuild the architecture, you can train it on your use case.
      3. janalsncm · · focus · HN ↗
        > it may raise some questions about the necessity of Jev's architecture

        When I hear “architecture” I am thinking number of parameters and latency.

        When I hear “accuracy” I think training recipe, data, and (later) number of parameters.

        So when you say that Jev’s architecture may not be necessary, the evidence I expect to see is comparable quality at comparable latency. Not equal quality at 2x latency and 4x the cost.

        1. kylecazar · · focus · HN ↗
          The post claims comparable quality at comparable latency (the difference in the latter only stemming from serving region).

          Hence my note about the risk of cost being the only moat.

      4. HarHarVeryFunny · · focus · HN ↗
        Whether Jev is something architecturally different from an LLM (bidirectional vs auto-regressive? different type of parallel readout head?) remains to be seen, or guessed, but LLMs are certainly fungible and are competing on price - they leapfrog each other from release to release, but overall they are all progressing in unison.

        When it comes to the high volume market of business automation, it seems that ultra-low cost, rather than expensive frontier intelligence, is exactly what you want, and low latency is also nice to have for customer-facing applications like customer service chatbots.

    3. kouteiheika · · focus · HN ↗
      > My question is why not use Jev instead? It's faster and cheaper.

      Because it's proprietary? By using an open weight model you're guaranteed that you can access it forever; if one provider bans you then you can go to another one (or you can self-host). With a proprietary, single-provider model locked behind an API if your access is revoked you're screwed.

    4. schainks · · focus · HN ↗
      Compliance. You can’t put Jev in a HIPAA compliant service, for example.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.