‹ BackHN Continuity

Thread

Kev: Tiny Jev-like family of decision models built on top of Qwen3.5

462 points · 211 comments · tosh

  1. monkeydust · · focus · HN ↗
    Bit of a Jev explosion going on. Is it because it's taking us back to a simpler time we understand better? Classification models have been around for a while.
    1. Tycho · · focus · HN ↗
      It’s because it’s practically useful and enabled things that were impractical previously.
      1. petesergeant · · focus · HN ↗
        > and enabled things that were impractical previously

        I think that there are not _that_ many use-cases that have been opened up by this that tool-calling on other models didn't solve already. Really depends what benchmark you're looking at. This one against BANKING77[0] has many issues, but suggests it's really not far off DeepSeek 4.1 Flash. This one against BoolQ[1] shows marginal improvement over Qwen3.6. This one against MMLU-Pro[2] (same author as the previous) shows significant improvements over two Qwen models.

        So there's definitely _some_ alpha there, but I don't think it's the sea-change that the hype would suggest; that is to say, yes, some things that weren't practical before are now, but many things were already very practical with the existing tools.

        0: <a href="https:&#x2F;&#x2F;sanand0.github.io&#x2F;llmevals&#x2F;jev&#x2F;" rel="nofollow">https:&#x2F;&#x2F;sanand0.github.io&#x2F;llmevals&#x2F;jev&#x2F;

        1: <a href="https:&#x2F;&#x2F;github.com&#x2F;ekzhang&#x2F;openjev-sglang&#x2F;blob&#x2F;a3554ed9e9c26d5d7b3b2184a524fc61779dbcc1&#x2F;evals&#x2F;results&#x2F;boolq-2026-09-18&#x2F;comparison&#x2F;report.md" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;ekzhang&#x2F;openjev-sglang&#x2F;blob&#x2F;a3554ed9e9c26...

        2: <a href="https:&#x2F;&#x2F;github.com&#x2F;ekzhang&#x2F;openjev-sglang&#x2F;blob&#x2F;a3554ed9e9c26d5d7b3b2184a524fc61779dbcc1&#x2F;evals&#x2F;results&#x2F;qwen38-27b&#x2F;report.md" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;ekzhang&#x2F;openjev-sglang&#x2F;blob&#x2F;a3554ed9e9c26...

        1. Fabricio20 · · focus · HN ↗
          The part about &quot;tool-calling on other models didn&#x27;t solve already&quot; is what gets you, sure I could tool call deepseek, glm or any other model, but the latency is huge and you get no confidence score. I gave JEV a shot via OpenRouter and it has a reply in less than 400ms, it&#x27;s fast enough and cheap enough that you can hook it up to a game loop for example (so highly state dependant) and it can do decisions in real time.
          1. petesergeant · · focus · HN ↗
            Please refer to the latency and price figures for DeepSeek V4.1 Flash in the first link I shared.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.