‹ BackHN Continuity

Thread

Clef: Open-weight decision models, and new RL fine-tuning platform

637 points · 217 comments · jasondavies

  1. manlymuppet · · focus · HN ↗
    Am I hearing this right, that they made a decision model based on Typesafe's new paradigm, and actually made a model better than Jev based on Typesafe's own ranking?

    And it's only been a few weeks.

    1. segmondy · · focus · HN ↗
      A lot of people claim to have made better than jev, there's a jev benchmark, I have tried many of those models and they eventually end up failing, a non trivial task which doesn't seem like much but reminds me of the svg pelican bench is games, have one of these decision/classifier models play a game, hook it up to the input, most of the ones that are supposedly on jev level end up playing a terrible game, showing that they are very narrow. Cloudflare doesn't compare to the top open bench alternatives, I just finished downloading it and will compare it to jev for non trivial tasks tonight.
      1. SebastianSosa · · focus · HN ↗
        Public benchmarks are easy to cheat, if I am typesafe I would also release a public benchmark to distract otherwise competent people in overfitting to a benchmark instead of making something actually useful. Diogo very much is against public benchmarks ;)
        1. Foobar8568 · · focus · HN ↗
          You take Qwen3.6 35b on a 5090rtx, and here you get a higher score than Jev, for 2sec more latency on average. So yeah it's not subsecond, but I am sure that if I had VC money, I could too get within 500ms too!
      2. verdverm · · focus · HN ↗
        watching Jev play Pokemon demonstrated this too, more hype than meat

        - I'd like a potion, are you sure, no, repeat

        - in and out of doors on loop

        - sisyphean effort in the cave

        - jev-ish level grinding

        It was impressive, beat pokemon for less than $2, but not all that interesting. People asking how different Math.Random plays pokemon would be, and at the other end, regular llms playing games.

      3. indoor47 · · focus · HN ↗
        Well, it makes sense:

        "Clef builds upon this concept, but uses a different base model as the backbone. We currently use Qwen as the base model and post-trained it to suit decision model use cases. "

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.