‹ BackHN Continuity

Thread

Kev: Tiny Jev-like family of decision models built on top of Qwen3.5

462 points · 211 comments · tosh

  1. hbarka · · focus · HN ↗
    If Jev is fundamentally trained using RLCD while you’re building on a Qwen model that was trained using RLHF, how can the resulting model be considered Jev-like?
    1. mohsen1 · · focus · HN ↗
      I can't find it but saw that if you give Jev English alphabet as choices and ask it in a loop what model it is, it would say Qwen

      also tried myself: <a href="https:&#x2F;&#x2F;console.typesafe.ai&#x2F;playground?share=shr_1690a3160f19c2e4851b6a2883bcf17a790" rel="nofollow">https:&#x2F;&#x2F;console.typesafe.ai&#x2F;playground?share=shr_1690a3160f1...

      1. c7b · · focus · HN ↗
        Once the first letter is Q, the rest is probably pretty determined. Can you see the confidence for the first letter (don&#x27;t want to accept the ToS to follow your link)?
        1. HenryMulligan · · focus · HN ↗
          I agree that once &quot;Q&quot; is selected, &quot;Qwen&quot; is by far the most likely choice. What I don&#x27;t get is why it would start by picking &quot;Q&quot;, one of the least-used letters in English, unless it already decided to say &quot;Qwen&quot;. Now, as others have pointed out, saying &quot;Qwen&quot; and being Qwen are two separate things (though I don&#x27;t get why they don&#x27;t just filter model declarations out of the dataset, or carefully replace them with theirs, as that would easily bias the model to always say their name).
          1. c7b · · focus · HN ↗
            Technically, Q is picked because it has the highest probability of all letters. But it makes a difference whether the probability for Q is barely above a uniform 1&#x2F;26~3.8% or whether that one letter concentrates &gt;50%. What I remember from reading the docs is that Jev gives you the full probabilities (and the confidence, which is something like normalized entropy).

            But in general, we might be reading too much into this. If I were to build something like this, a Qwen model would be among the first things I&#x27;d reach for too. Initially just prompted inside a little harness to guarantee you get the desired output. Next step would be finetuning, finally training your own foundation model, if you can muster the funding. In this fast-moving space, I think it&#x27;s quite understandable that they&#x27;d go public with an MVP asap, so likely not much training on their own. And even if they&#x27;re finetuning, Qwen&#x27;s baked-in answer (through Alibaba&#x27;s finetuning) seems likely to survive unless it was explicitly overridden.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.