‹ BackHN Continuity

Thread

Clef: Open-weight decision models, and new RL fine-tuning platform

637 points · 217 comments · jasondavies

  1. amluto · · focus · HN ↗
    I’ll go out on a limb and suggest that I don’t think a Jev-like model is particularly useful unless you can fine tune it. The Jev API has zero ability to pass in a prior [0], and, if you can neither pass in a prior nor fine tune for your system, you will get an output that may be almost meaningless.

    I’d love to see someone build a model of this sort that can actually accept priors and do something intelligent with them.

    [0] You can feed Jev a prior as text. I’ve tried it. It works poorly.

    1. brokensegue · · focus · HN ↗
      I think better than priors would be a closed loop where you tell it what the right answer was (or some signal) and they monitor and fine-tune for you
    2. okpatil · · focus · HN ↗
      We were able to completely automate 20,100 token prompts with At0m[<a href="https:&#x2F;&#x2F;at0m.pienomial.com&#x2F;" rel="nofollow">https:&#x2F;&#x2F;at0m.pienomial.com&#x2F;].

      We believe entire compliance workflows (even multilingual) could be automated.

      Would you like to get a demo ?

      1. sheepscreek · · focus · HN ↗
        You’re coming on a bit strongly - a couple of comments with a link is sufficient. Before trying to sell, try to genuinely further the conversation, provide some useful knowledge in return for the reader’s attention.
        1. okpatil · · focus · HN ↗
          Point taken. Let me explain if you allow me.

          It is possible with deterministic decision models, such as At0m, to gauge the probabilities at every decision. This behavior in addition to hard coded logic, it is possible to completely replicate a prompt&#x27;s logic.

          Using Fable 5.1, it is a matter of minutes.

          I believe that most of the compliance check documents will be a solved problem, 3-6 months in future.

          None of the LLMs can do it.

          Hence I asked to the comment poster if he would want to demo, so that I can show it to him, how to do it step by step. By bad, if it came out too strongly.

    3. sheepscreek · · focus · HN ↗
      Also one of the more interesting features of Jev is the confidence rating that hardly any Jev-cc talks about.
      1. amluto · · focus · HN ↗
        It seems interesting to me only in the sense of being useless. From the horse’s mouth:

        &gt; Confidence is derived from the probabilities

        <a href="https:&#x2F;&#x2F;docs.typesafe.ai&#x2F;confidence" rel="nofollow">https:&#x2F;&#x2F;docs.typesafe.ai&#x2F;confidence

        (Why is it much easier to find AI-slop websites quoting this than it is to find the actual documentation?)

        My inner Bayesian would like for Jev to provide something resembling “evidence”, although I admit that one might ask Jev questions that are somewhat awkward to treat as typical Bayesian questions. If I ask “will this PR be merged”, it’s kind of strange to contemplate the probability of a PR conditioned in that PR being merged in the future. But I bet there is a way to formalize a prior-free classifier in a way that makes Bayesians and non-Bayesians happy, possibly involving actual learned probabilities and confidence levels. If you read the literature on scoring rules, you will find that classifier scores do somewhat naturally decompose into a few interpretable terms.

    4. mikeocool · · focus · HN ↗
      It seems like jev&#x27;s major advantage over existing classifiers is that I dont have train it.

      If I have to gather and tag data to fine-tune Jev, I can probably just train an &quot;old school&quot; classifier model and make it even cheaper, faster, and just as accurate.

    5. jdthedisciple · · focus · HN ↗
      I suppose a sort of prior-proxy can be encapsulated by a carefully written system prompt.
      1. amluto · · focus · HN ↗
        The straightforward approach of literally staying a numeric prior has some effect but not the correct effect.
    6. AnthusAI · · focus · HN ↗
      You can&#x27;t fine-tune Jev itself but you can train an ML model that uses outputs from Jev as inputs. Which does enable you to &#x27;fine-tune&#x27; your overall model.

      You can also improve your Jev classifications based on your ongoing data if you&#x27;re labeling it continuously, especially if you&#x27;re explaining the reasoning in the feedback labels. You can identify new elements of the rubric and add them to the list of classifications that Jev produces, and then those become new features for your ML model.

      Two levers of control for using data to make a Jev-based classifier model continuously better-aligned.

      1. AnthusAI · · focus · HN ↗
        Example of that: <a href="https:&#x2F;&#x2F;anth.us&#x2F;blog&#x2F;fine-tuning-jev&#x2F;" rel="nofollow">https:&#x2F;&#x2F;anth.us&#x2F;blog&#x2F;fine-tuning-jev&#x2F;
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.