‹ BackHN Continuity

Thread

Clef: Open-weight decision models, and new RL fine-tuning platform

637 points · 217 comments · jasondavies

  1. manlymuppet · · focus · HN ↗
    Am I hearing this right, that they made a decision model based on Typesafe's new paradigm, and actually made a model better than Jev based on Typesafe's own ranking?

    And it's only been a few weeks.

    1. TeMPOraL · · focus · HN ↗
      It's not a "new paradigm", it's a low-hanging fruit that's been lying around for years; Typesafe were the first to bother to stop and pick it up, and market the shit out of it. But it was still a low-hanging fruit.

      There are many, many of those left around, because AI frontier is moving forward so fast, everyone is racing ahead. Which is why I laugh when people say AI is not transformative and LLMs are a dead end (and my favorite, "what are we going to do with all those GPUs when the bubble pops?"). Even if SOTA LLMs hit a hard capability limit tomorrow and never advanced again, there's a good decade of growth and advancement to be extracted just from all the low-hanging fruits that were left unpicked along the way.

      1. seizethecheese · · focus · HN ↗
        Name a few of these low hanging fruit left around.
        1. TeMPOraL · · focus · HN ↗
          Jev is one.

          Diffusion transformers are not "easy" but underfunded.

          Random one in terms of applications: getting GPT-4-level[0] LLMs to operate at hundreds of tokens per second on edge hardware - opens up so many possibilities I'm probably unable to imagine half of them.

          E.g. Imagine spellcheck/predictive text (or code autocomplete) where the model is able to process a whole paragraph + surrounding application/system context in between keystrokes. Or an OS being able to reliably guess what you're doing in real-time, in between your UI interactions, and offer actually helpful contextual reactions.

          Or imagine finally funding some decent studies into exploring the models as computational artifacts - studying their latent spaces, how they form and how they model reality internally.

          Or imagine automated sliding doors that don't suck.

          --

          [0] - Or anything substantially better than BERT-level models used in Jev or that demo from the company doing inference ASICs, that has a chatbot online that does 14 kilotokens per second.

          1. mdp2021 · · focus · HN ↗
            There are "low hanging fruits" - easier to achieve goals -, and there are super-fruits, milestone-fruits.

            Among the most important ones:

            -- the long-known Problem of Transparency, applied to the apparent emergent intelligence in NNs. Why does it happen - in detail?

            -- then, a Theory of Apparent Intelligence through NNs. Transforming the results achieved into a Science. Which allows to do what we are doing - but in a lean and targeted way.

            -- then, a General Theory of Intelligence, that includes the above to go beyond current architectures and get those features of Intelligence we expect and still not have.

            The long-term direction we got into must lead to this.

            (You note a ponderant detail of the above when you note the importance of explaining the emergence of a World Model from a Language Model.)

            1. TeMPOraL · · focus · HN ↗
              Those are the absolutely fascinating parts, and I sincerely hope AI won't get out of control before we're able to tackle some of these.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.