‹ BackHN Continuity

Thread

Clef: Open-weight decision models, and new RL fine-tuning platform

637 points · 217 comments · jasondavies

  1. dgacmu · · focus · HN ↗
    Since it supports vision, I tested it (an 8 bit quant) against one of my pet problems of coin classification (US coins, but not beauty shots of them, just lots of cell phone images on a piece of paper). It did .. horribly. About 41% accuracy. My local gemma4:27b gets 53% and runs 4x faster. Not too surprising but seemed worth poking at it.
    1. leopoldj · · focus · HN ↗
      The model supports both instruction and RL fine-tuning. It looks like people have started training it already [1]. The RL tuning may be only available to the Cloudflare customers, I couldn't be sure.

      1. <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;solanaclawd&#x2F;clef-solana-research" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;solanaclawd&#x2F;clef-solana-research

      1. ActivePattern · · focus · HN ↗
        The biggest upside to using a zero-shot classification model is that it&#x27;s actually zero-shot, i.e., you don&#x27;t need to construct task-specific training sets.

        If it needs to be fine-tuned, then why not fine-tune smaller and cheaper models that are only as big as the task requires?

        1. leopoldj · · focus · HN ↗
          I hear what you&#x27;re saying. But this is not such a model. The article says that various internal teams had to fine-tune Clef to suit their specific needs. They&#x27;re even rolling out a RL training product line for it.
          1. dgacmu · · focus · HN ↗
            Fair. Though in this case my actual classifier (dino v3 features fed into a very shallow DNN) works surprisingly well and is insanely faster than using an LLM -- so it really was the one shot performance I was hoping for.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.