‹ BackHN Continuity

Thread

Clef: Open-weight decision models, and new RL fine-tuning platform

637 points · 217 comments · jasondavies

  1. agrippanux · · focus · HN ↗
    I'm a big fan of Cloudflare products. I recently stuck Jev in front of a Cloudflare-hosted Ollama model for chat/username moderation, so I was excited to test out Clef. The setup is user send a message -> Jev does first pass to see if it's toxic/hate speech/profane, if Jev is unsure then Ollama on Workers AI takes a deeper look.

    Clef was 2-3x slower and worse (it caught less hate speech) than Jev. Overall disappointing.

    1. wongarsu · · focus · HN ↗
      It's likely no coindicence that they only show "median latency", not how latency scales with input size.

      For small inputs Jev is slow, but its latency curve is very flat. Fine-tuning a decently-sized llm (like this 27B model) gives you something that's faster on small input sizes, but even with moderate contexts quickly becomes much slower than Jev. Characterizing it as "faster than Jev" is very misleading, unless you know all your questions have tiny context (less than 1k tokens or so)

      1. tomrod · · focus · HN ↗
        True. But do I understand that Clef is multimodal and Jev is only text?
        1. wongarsu · · focus · HN ↗
          yes. Jev is text-only, while Clef and a lot of the other alternatives are fine-tunes of multi-modal models, so you get image-input basically for free. Actual decision quality based on images is a bit up in the air though, I am not aware of any benchmarks testing that
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.