‹ BackHN Continuity

Thread

Clef: Open-weight decision models, and new RL fine-tuning platform

637 points · 217 comments · jasondavies

  1. manlymuppet · · focus · HN ↗
    Am I hearing this right, that they made a decision model based on Typesafe's new paradigm, and actually made a model better than Jev based on Typesafe's own ranking?

    And it's only been a few weeks.

    1. TeMPOraL · · focus · HN ↗
      It's not a "new paradigm", it's a low-hanging fruit that's been lying around for years; Typesafe were the first to bother to stop and pick it up, and market the shit out of it. But it was still a low-hanging fruit.

      There are many, many of those left around, because AI frontier is moving forward so fast, everyone is racing ahead. Which is why I laugh when people say AI is not transformative and LLMs are a dead end (and my favorite, "what are we going to do with all those GPUs when the bubble pops?"). Even if SOTA LLMs hit a hard capability limit tomorrow and never advanced again, there's a good decade of growth and advancement to be extracted just from all the low-hanging fruits that were left unpicked along the way.

      1. seizethecheese · · focus · HN ↗
        Name a few of these low hanging fruit left around.
        1. murkt · · focus · HN ↗
          Easy to reach doesn’t automatically mean “easy to see”.
        2. [deleted] · · focus · HN ↗

          [deleted]

        3. TeMPOraL · · focus · HN ↗
          Jev is one.

          Diffusion transformers are not "easy" but underfunded.

          Random one in terms of applications: getting GPT-4-level[0] LLMs to operate at hundreds of tokens per second on edge hardware - opens up so many possibilities I'm probably unable to imagine half of them.

          E.g. Imagine spellcheck/predictive text (or code autocomplete) where the model is able to process a whole paragraph + surrounding application/system context in between keystrokes. Or an OS being able to reliably guess what you're doing in real-time, in between your UI interactions, and offer actually helpful contextual reactions.

          Or imagine finally funding some decent studies into exploring the models as computational artifacts - studying their latent spaces, how they form and how they model reality internally.

          Or imagine automated sliding doors that don't suck.

          --

          [0] - Or anything substantially better than BERT-level models used in Jev or that demo from the company doing inference ASICs, that has a chatbot online that does 14 kilotokens per second.

          1. aeve890 · · focus · HN ↗
            >Random one in terms of applications: getting GPT-4-level[0] LLMs to operate at hundreds of tokens per second on edge hardware

            That's low hanging for you?

            1. blurbleblurble · · focus · HN ↗
              It's likely quite close. There are so many papers proving concepts that would bring this, they just haven't been combined in production.
            2. msdz · · focus · HN ↗
              Maybe they meant in the sense of “untapped potential”, because so far a lot of the focus has been on increasing model capabilities, not necessarily performance/power budget.
              1. TeMPOraL · · focus · HN ↗
                Yes. Point is, it's untapped only because everyone is running in the race (even if out of curiosity), and there's just not enough people with means to tap into these side threads. For the past few years, there's been many interesting papers that circulated the industry, got recognized as worthwhile pursuits, and then dropped because running behind the Big Vendors had massively better ROI.
            3. ekabod · · focus · HN ↗
              That's a high hanging fruit, not low.
            4. guyomes · · focus · HN ↗
              If we throw in hardware dedicated to a specific LLM, it seems to be a rather low hanging fruit. Especially considering that this is already happening for vision models [1].

              [1]: &quot;FPGA-based CNN Acceleration using Pattern-Aware Pruning&quot; <a href="https:&#x2F;&#x2F;inria.hal.science&#x2F;hal-04689673&#x2F;document" rel="nofollow">https:&#x2F;&#x2F;inria.hal.science&#x2F;hal-04689673&#x2F;document

              1. mdp2021 · · focus · HN ↗
                &gt; hardware dedicated to a specific LLM

                That wording screams &quot;Taalas&quot;. Which, importantly, is not the only player trying to abate the distance between data and arithmetics...

            5. TeMPOraL · · focus · HN ↗
              Yes. It&#x27;s well within realm of possibility, but so far wasn&#x27;t pursued because the Big Vendors went all-in into capability growth (rightfully testing &quot;the bitter lesson&quot; to its limits) and got themselves stuck in an arms race, while everyone else is barely keeping up and&#x2F;or starstruck with fascination, exploring what these models can do.

              This got everyone racing forward and right now there is not enough human attention left in the world to productionize this, or any of the other &quot;side threads&quot;. When the race slows down, people will catch up, branch out, and loop back.

              1. blurbleblurble · · focus · HN ↗
                Just like renewable energy and so many other things. Hyperconcentration of capital is really tragic. I hope things turn around.
                1. TeMPOraL · · focus · HN ↗
                  They will. That&#x27;s the fallacy of the &quot;S-curve&quot; everyone likes to commit these days actually giving a positive outlook.

                  Assuming it won&#x27;t get to full RSI, the current approach will burn out - most likely economically. The race slows down, people branch out, look back, start picking up the &quot;untapped potential&quot;&#x2F;low-hanging fruits, and you have new S-curves launching in place of the one that just tapered off (hence a fallacy - a stack of S-curves adds up to continuing exponential growth).

                  In other words: it comes and goes. Hyperconcentrated capital will eventually deconcentrate.

            6. mdp2021 · · focus · HN ↗
              &gt; That&#x27;s low hanging for you

              An important part of the industry is studying that: it is built-up effort. Sooner or later, the fruits will be harvested. The targeted preparation has been there for years now.

            7. Twirrim · · focus · HN ↗
              We already have examples of LLMs running 16k+ tokens a second using custom ASICs.

              It&#x27;s down at the moment (Not sure if it&#x27;ll return?) but Chat Jimmy[0] produced by Taalas[1] was powered by an ASIC running Llama 3.1 8B, and hitting 17,000 tokens&#x2F;sec. It was amazing to use, you&#x27;d no sooner have hit enter than you had a full response back. I actually found its speed to be a problem for interactive stuff, as every answer got several paragraphs I&#x27;d then wade through, vs a model populating text closer to my reading speed.

              I appreciate, there are differences between an 8 billion parameter model and something GPT-4-ish, but we&#x27;re currently in the middle of a race between a half dozen or so companies to produce the next best frontier model, which requires their infrastructure to be dynamic.

              We really don&#x27;t always need newer better faster stronger models, there&#x27;s quite a lot of room for &quot;good enough&quot; where getting 17kt&#x2F;s at significantly lower power would be amazing.

              [0] <a href="https:&#x2F;&#x2F;chatjimmy.ai&#x2F;" rel="nofollow">https:&#x2F;&#x2F;chatjimmy.ai&#x2F; [1] <a href="https:&#x2F;&#x2F;taalas.com&#x2F;" rel="nofollow">https:&#x2F;&#x2F;taalas.com&#x2F;

              1. AshamedBadger56 · · focus · HN ↗
                &gt;I actually found its speed to be a problem for interactive stuff, as every answer got several paragraphs I&#x27;d then wade through, vs a model populating text closer to my reading speed.

                It would be interesting to pair the super fast model with a normal speed model. Have the super fast one do all the background research, code writing, etc. The normal model would just relay the needed info to you at a more reasonable pace.

                1. drewstiff · · focus · HN ↗
                  Why do you even need a second model for that? There are many simpler ways to slow down the rate that the near-immediate response is pseudo-typed onto the screen.
          2. flipping_beacon · · focus · HN ↗
            Definitely agree with edge computation, although inference extensively researched and funded if SOTA LLMs hit a dead end tomorrow,there is still a lot to explore and research in inference and edge computation
          3. blurbleblurble · · focus · HN ↗
            Diffusion models combined with these new looping techniques are gonna change the whole conversation about efficiency. Imagine control net but in one or more conceptual latent spaces.

            But also harnesses and more generally new insights on &quot;the control flow problem&quot; could end up squeezing a ton of performance out of small models.

          4. Amekedl · · focus · HN ↗
            yeah your reply, nobody can predict the future.

            Enough stuff can happen, software use itself might change, and that could really cause anything. &quot;What will we do with all the gpus&quot; might become a question if for a magnitude of tech and reasons leaked-opus-9 runs on a macbook m6 or 7

          5. seizethecheese · · focus · HN ↗
            I commend you for actually answering, independent of what I think of the answers.
            1. dominotw · · focus · HN ↗
              not a very good answer though
              1. tomrod · · focus · HN ↗
                Yes yes, most comments and responses fail to contribute to conversation (though I disagree with your particular assessment); it thus remains important to continue to share our own thoughts as we battle through thesis and antithesis to arrive at synthesis.
          6. alightsoul · · focus · HN ↗

            [dead]

          7. mdp2021 · · focus · HN ↗
            There are &quot;low hanging fruits&quot; - easier to achieve goals -, and there are super-fruits, milestone-fruits.

            Among the most important ones:

            -- the long-known Problem of Transparency, applied to the apparent emergent intelligence in NNs. Why does it happen - in detail?

            -- then, a Theory of Apparent Intelligence through NNs. Transforming the results achieved into a Science. Which allows to do what we are doing - but in a lean and targeted way.

            -- then, a General Theory of Intelligence, that includes the above to go beyond current architectures and get those features of Intelligence we expect and still not have.

            The long-term direction we got into must lead to this.

            (You note a ponderant detail of the above when you note the importance of explaining the emergence of a World Model from a Language Model.)

            1. TeMPOraL · · focus · HN ↗
              Those are the absolutely fascinating parts, and I sincerely hope AI won&#x27;t get out of control before we&#x27;re able to tackle some of these.
            2. patcon · · focus · HN ↗
              If anyone is interested, following Dr Michael Levin&#x27;s Thoughtforms.life podcast is the cutting edge of where all this previously fuzzy stuff is becoming more concrete. So long as you can tolerate distinguished scientists flailing about as they discuss consciousness and life and developmental biology (and other less-obviously living things, like algorithms) as involving &quot;free lunches&quot; and &quot;ingressing patterns from the platonic realm&quot; :)
        4. hobofan · · focus · HN ↗
          Closely connected to decision models: A good library to do ranking based on pairwise ranking on multiple attributes. By using a decision model (especially one that can make decisions on multiple fields at the same time) this becomes a lot faster and more powerful. Could make for a pretty nice search reranker as well as prioritizer for many problems.

          Of course you can also do ranking one-off with a decision model, but this likely less stable, and by doing pairwise ranking you can also relatively quickly do incremental inserts to the list.

        5. sarkarghya · · focus · HN ↗
          I can imagine advancements on making smaller models work together better instead of a generalized core. Imagine a community or city having a <a href="https:&#x2F;&#x2F;pirateface.co&#x2F;" rel="nofollow">https:&#x2F;&#x2F;pirateface.co&#x2F; so that the shard of the model that you need can be streamed in with minimal latency with your box only holding the minimal version (say deepseek v4 flash as orchestrator) of the model that you use on day to day basis.

          We have overcome split brain problems before so this wont be our first

          1. esseph · · focus · HN ↗
            You&#x27;re putting a lot of trust in uncorrupted, untrusted, unknown models (potentially).
        6. CamperBob2 · · focus · HN ↗
          An example I like to use is: compare the quality and scope of games released with a brand-new console to the ones released for that console towards the end of its life, when everyone has learned how to take advantage of whatever weird, wacky hardware Sony invented for that console generation.

          There is still a lot we don&#x27;t know about how to get the most out of existing LLM components from a speed or cognitive-performance perspective. People could easily spend the next decade studying and refining what&#x27;s been built so far, even if no new, original approaches ever arrive.

          1. randomNumber7 · · focus · HN ↗
            I never enjoyed playing ps4 when it is constantly louder than my vacuum cleaner.
        7. sroussey · · focus · HN ↗
          Just look at all the model type on hugging face. LLMs are a small percentage.
      2. btown · · focus · HN ↗
        Synthetic data is key here! Compared to 5 years ago, we now have oracles that can generate massive data sets of perfectly labeled multimodal data, practically for free. The number of architectures that can benefit from that is innumerable, and far beyond just LLMs themselves. On top of this, LLMs can implement any architectural ideas you have, and write custom tools to manage training and evaluation.

        Whether or not LLMs can self-improve their frontier capabilities, they can absolutely create a wake for themselves that accelerates everything else that&#x27;s training on their synthetic data. We&#x27;ll see every architecture of the past 40 years suddenly show leaps and bounds.

        1. gutchapa · · focus · HN ↗
          Agreed on the directions, it might appear to be piece of cake, but the &quot;practically for free&quot; part hides the hard bit: realism. Try simulating flight-search data with source, destination, flight numbers, routes, schedules , try a diff it against a real corpus. Even frontier models produce data that looks right at first sight, but breaks on joint distributions (a flight number that never flies that route, layovers that violate minimum connect times). last mile is the daunting task.
          1. btown · · focus · HN ↗
            To lean into your example: if you want to design travel search, you might envision a system that goes from a user&#x27;s typed query -&gt; a structured query -&gt; executed against a database of real-time flights -&gt; judging the result for feasibility of the resulting proposals created by those queries.

            What you can do with LLMs is reverse this: from any numbers of snapshots of flight data, you can create large numbers of plausible user queries, based on your data, that are known to be feasible or infeasible. And now you have a labeled data set to train a model that focuses solely on the query-creation and judgment systems. And you can experiment with whether having a more flexible query protocol leads to higher success rates without sacrificing accuracy, or whether you can generate that last-mile feasibility check as a combination of auditable code checks alongside AI-based judgment.

            LLMs don&#x27;t absolve you of having to break down your system architectures into components that have well-defined boundaries (though certainly they can help with that design). They do make those components feasible to solve at scale.

      3. dzonga · · focus · HN ↗
        my take there&#x27;s a lot of low hanging fruit in applying small models to knowledge economy workflows.

        then vision &amp; robotics.

        while everyone&#x27;s chasing the frontier.

      4. outofpaper · · focus · HN ↗
        Yup, one output token, and read the logprobs of the posible tokens that could have been thos 1st token. Plenty of systems already do this; Typesafe&#x27;s fundraising and marketing just made it visible.

        Good to see interest broadening beyond &quot;just extend thinking.&quot; More approaches in the toolbox means fewer problems get treated as nails.

      5. tomrod · · focus · HN ↗
        &gt; what are we going to do with all those GPUs when the bubble pops?

        Your comment here made me laugh, because I think we will finally be able to play Crysis at 10fps.

        Just kidding, of course. GPU half lives are quite a bit less than standard compute half lives, no?[0] That&#x27;s what I&#x27;ve been trying to understand regarding data centers focusing as GPU clusters -- seems like the ROI window would have to be very short for the capitalization.

        [0] <a href="https:&#x2F;&#x2F;www.tomshardware.com&#x2F;pc-components&#x2F;gpus&#x2F;datacenter-gpu-service-life-can-be-surprisingly-short-only-one-to-three-years-is-expected-according-to-unnamed-google-architect" rel="nofollow">https:&#x2F;&#x2F;www.tomshardware.com&#x2F;pc-components&#x2F;gpus&#x2F;datacenter-g...

    2. slopnt · · focus · HN ↗
      They have to have decision models already in production. Part of their business is detecting bots, DDoSers and spammers.
      1. alightsoul · · focus · HN ↗
        Yeah that&#x27;s a decision tree, random Forest or some other machine learning classifier. They have existed for a long time
        1. smallmancontrov · · focus · HN ↗
          I&#x27;m all for rebranding discriminative models as decision models, though.

          &quot;Discriminative&quot; always had pointlessly bad optics, but I knew it was over when I started seeing prominent machine learning researchers who p=100% knew better describe discriminative models as generative because that was the buzzword of the year. &quot;Decision model&quot; sells the value proposition much better and doesn&#x27;t sound like an anti-woke crusade.

          1. tomrod · · focus · HN ↗
            Decision models have long been called `choice` models to may understanding. Huge statistical literature on dynamic discrete choice models and inference!
        2. eastdakota · · focus · HN ↗
          I think you just called me old.
    3. segmondy · · focus · HN ↗
      A lot of people claim to have made better than jev, there&#x27;s a jev benchmark, I have tried many of those models and they eventually end up failing, a non trivial task which doesn&#x27;t seem like much but reminds me of the svg pelican bench is games, have one of these decision&#x2F;classifier models play a game, hook it up to the input, most of the ones that are supposedly on jev level end up playing a terrible game, showing that they are very narrow. Cloudflare doesn&#x27;t compare to the top open bench alternatives, I just finished downloading it and will compare it to jev for non trivial tasks tonight.
      1. SebastianSosa · · focus · HN ↗
        Public benchmarks are easy to cheat, if I am typesafe I would also release a public benchmark to distract otherwise competent people in overfitting to a benchmark instead of making something actually useful. Diogo very much is against public benchmarks ;)
        1. Foobar8568 · · focus · HN ↗
          You take Qwen3.6 35b on a 5090rtx, and here you get a higher score than Jev, for 2sec more latency on average. So yeah it&#x27;s not subsecond, but I am sure that if I had VC money, I could too get within 500ms too!
      2. verdverm · · focus · HN ↗
        watching Jev play Pokemon demonstrated this too, more hype than meat

        - I&#x27;d like a potion, are you sure, no, repeat

        - in and out of doors on loop

        - sisyphean effort in the cave

        - jev-ish level grinding

        It was impressive, beat pokemon for less than $2, but not all that interesting. People asking how different Math.Random plays pokemon would be, and at the other end, regular llms playing games.

      3. indoor47 · · focus · HN ↗
        Well, it makes sense:

        &quot;Clef builds upon this concept, but uses a different base model as the backbone. We currently use Qwen as the base model and post-trained it to suit decision model use cases. &quot;

    4. heliosAtwork · · focus · HN ↗
      There was some parallel independent work from Sep 28. They mention it in the blog post. But I am sure Jev has opened a lot of eyes on the possibilities.

      &quot;In the same week that Jev came out, we posted about some experiments [1] we had with our own homegrown decision model.&quot;

      [1] <a href="https:&#x2F;&#x2F;x.com&#x2F;michellechen&#x2F;status&#x2F;2101091012559151480" rel="nofollow">https:&#x2F;&#x2F;x.com&#x2F;michellechen&#x2F;status&#x2F;2101091012559151480

    5. lofaszvanitt · · focus · HN ↗
      Race to the bottom. The one with the most resources wins.
    6. fwip · · focus · HN ↗
      From their blog post, it sounds like it is still 10x as slow as Jev.

      As far as I&#x27;ve seen, all of the Jev-compatible projects simply take an LLM, hack off a layer or two at the end, and call it good. Some of them spend more work than others trying to back-estimate in accurate probabilities.

      1. rahimnathwani · · focus · HN ↗
        Yeah I think people were already using either of these as part of traditional software workflows):

        A) Structured outputs from LLMs (doesn&#x27;t need fine tuning but can be expensive)

        B) Classification output from fine-tuned BERT-like or GLiNER models (is calibrated well and has cheap&#x2F;fast inference)

        What Jev did is combine the advantages of both A and B into one model&#x2F;product, and create a really good API.

        They claim that a key innovation is how they&#x27;ve trained the model using what they call RLCD (RL from calibrated decisions). So it&#x27;s not just that you can get the outputs (which is easy to add to any LLM) but that the different primitives they expose (Choice, Score, Noul) have each been calibrated. For example, they claim that if you use the Score primitive (which gives you probabilities along a bunch of choices representing a continuum) that&#x27;s not just using the more general &#x27;Choice&#x27; primitive under the hood. It&#x27;s been calibrated separately.

        I don&#x27;t know how many of the Jev-like things we&#x27;ve seen do that. For example Cloudflare offers a Jev-like model with the same API, and which they say was trained with RLCD. But I don&#x27;t know whether Choice and Score are different under the hood, or whether Score is just sugar on top of Choice. (Should be easy to test this, but I haven&#x27;t done it.)

    7. nater5000 · · focus · HN ↗
      &gt;And it&#x27;s only been a few weeks.

      You make it sound like this is some noteworthy timeframe. I&#x27;d be more surprised if Jev wasn&#x27;t immediately made obsolete within days of release (which was effectively the case), especially when backed by a company like Cloudflare lol

      ML moves quick. Add the extreme hype and cash floating around in this space and you can expect that anything resembling something novel and relatively untapped is going to be pounced on and turned over basically immediately.

    8. xnx · · focus · HN ↗
      What&#x27;s the hype with Jev? Hasn&#x27;t Gemini had structured outputs since November 2025? <a href="https:&#x2F;&#x2F;ai.google.dev&#x2F;gemini-api&#x2F;docs&#x2F;structured-output" rel="nofollow">https:&#x2F;&#x2F;ai.google.dev&#x2F;gemini-api&#x2F;docs&#x2F;structured-output
      1. manlymuppet · · focus · HN ↗
        Traditional LLMs doing structured output is like trying to fit a square peg in a round hole. It&#x27;s just not the right tool for the job. They can do structured output, but it&#x27;s janky.

        With Jev, structured output is its native format. Jev is just way faster, cheaper, and outright better for a lot of things.

        Note that while all of this is great, this is nowhere near a sort of &quot;ChatGPT moment&quot;. It&#x27;s a cool new thing, and it&#x27;s way better at certain tasks, which could be big. That&#x27;s all though.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.