‹ BackHN Continuity

Thread

Clef: Open-weight decision models, and new RL fine-tuning platform

637 points · 217 comments · jasondavies

  1. pdlug · · focus · HN ↗
    I love what Cloudflare is doing generally so I was excited to try Clef in my evals on a real task vs Jev: should an agent's knowledge-base write go to human review?

    Quality: close (recall 0.98 vs 1.00) Hosted p50: Clef ~850ms, Jev ~110ms Clef-flash: over-escalates

    Data + script: <a href="https:&#x2F;&#x2F;github.com&#x2F;nicia-ai&#x2F;admission-decision-eval" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;nicia-ai&#x2F;admission-decision-eval

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.