‹ BackHN Continuity

Thread

Clef: Open-weight decision models, and new RL fine-tuning platform

637 points · 217 comments · jasondavies

  1. manlymuppet · · focus · HN ↗
    Am I hearing this right, that they made a decision model based on Typesafe's new paradigm, and actually made a model better than Jev based on Typesafe's own ranking?

    And it's only been a few weeks.

    1. TeMPOraL · · focus · HN ↗
      It's not a "new paradigm", it's a low-hanging fruit that's been lying around for years; Typesafe were the first to bother to stop and pick it up, and market the shit out of it. But it was still a low-hanging fruit.

      There are many, many of those left around, because AI frontier is moving forward so fast, everyone is racing ahead. Which is why I laugh when people say AI is not transformative and LLMs are a dead end (and my favorite, "what are we going to do with all those GPUs when the bubble pops?"). Even if SOTA LLMs hit a hard capability limit tomorrow and never advanced again, there's a good decade of growth and advancement to be extracted just from all the low-hanging fruits that were left unpicked along the way.

      1. btown · · focus · HN ↗
        Synthetic data is key here! Compared to 5 years ago, we now have oracles that can generate massive data sets of perfectly labeled multimodal data, practically for free. The number of architectures that can benefit from that is innumerable, and far beyond just LLMs themselves. On top of this, LLMs can implement any architectural ideas you have, and write custom tools to manage training and evaluation.

        Whether or not LLMs can self-improve their frontier capabilities, they can absolutely create a wake for themselves that accelerates everything else that's training on their synthetic data. We'll see every architecture of the past 40 years suddenly show leaps and bounds.

        1. gutchapa · · focus · HN ↗
          Agreed on the directions, it might appear to be piece of cake, but the "practically for free" part hides the hard bit: realism. Try simulating flight-search data with source, destination, flight numbers, routes, schedules , try a diff it against a real corpus. Even frontier models produce data that looks right at first sight, but breaks on joint distributions (a flight number that never flies that route, layovers that violate minimum connect times). last mile is the daunting task.
          1. btown · · focus · HN ↗
            To lean into your example: if you want to design travel search, you might envision a system that goes from a user's typed query -> a structured query -> executed against a database of real-time flights -> judging the result for feasibility of the resulting proposals created by those queries.

            What you can do with LLMs is reverse this: from any numbers of snapshots of flight data, you can create large numbers of plausible user queries, based on your data, that are known to be feasible or infeasible. And now you have a labeled data set to train a model that focuses solely on the query-creation and judgment systems. And you can experiment with whether having a more flexible query protocol leads to higher success rates without sacrificing accuracy, or whether you can generate that last-mile feasibility check as a combination of auditable code checks alongside AI-based judgment.

            LLMs don't absolve you of having to break down your system architectures into components that have well-defined boundaries (though certainly they can help with that design). They do make those components feasible to solve at scale.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.