‹ BackHN Continuity

Thread

Clef: Open-weight decision models, and new RL fine-tuning platform

637 points · 217 comments · jasondavies

  1. agrippanux · · focus · HN ↗
    I'm a big fan of Cloudflare products. I recently stuck Jev in front of a Cloudflare-hosted Ollama model for chat/username moderation, so I was excited to test out Clef. The setup is user send a message -> Jev does first pass to see if it's toxic/hate speech/profane, if Jev is unsure then Ollama on Workers AI takes a deeper look.

    Clef was 2-3x slower and worse (it caught less hate speech) than Jev. Overall disappointing.

    1. teleforce · · focus · HN ↗
      In "How we trained Clef" section Cloudflare mentioned that they initially based their Jev-like system on DiffussionGemma (DJev) model but then changed to Qwen with post training for Clef, but never mentioned any reason and justification for the change.

      Although they included the benchmark against DJev and Clef is better, perhaps if you can test it to see the real-world performance.

      1. nostrebored · · focus · HN ↗
        Diffusion Gemma sucks to train if you are not already mostly in distribution. The reason diffusion Gemma is fast is a learned denoising that balances quality and speed. So the further your data is from what already happens, the slower it gets. In almost every case I’ve tried, you might as well just train Nemotron.
    2. wongarsu · · focus · HN ↗
      It's likely no coindicence that they only show "median latency", not how latency scales with input size.

      For small inputs Jev is slow, but its latency curve is very flat. Fine-tuning a decently-sized llm (like this 27B model) gives you something that's faster on small input sizes, but even with moderate contexts quickly becomes much slower than Jev. Characterizing it as "faster than Jev" is very misleading, unless you know all your questions have tiny context (less than 1k tokens or so)

      1. tomrod · · focus · HN ↗
        True. But do I understand that Clef is multimodal and Jev is only text?
        1. wongarsu · · focus · HN ↗
          yes. Jev is text-only, while Clef and a lot of the other alternatives are fine-tunes of multi-modal models, so you get image-input basically for free. Actual decision quality based on images is a bit up in the air though, I am not aware of any benchmarks testing that
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.