‹ BackHN Continuity

Thread

Clef: Open-weight decision models, and new RL fine-tuning platform

637 points · 217 comments · jasondavies

  1. manlymuppet · · focus · HN ↗
    Am I hearing this right, that they made a decision model based on Typesafe's new paradigm, and actually made a model better than Jev based on Typesafe's own ranking?

    And it's only been a few weeks.

    1. TeMPOraL · · focus · HN ↗
      It's not a "new paradigm", it's a low-hanging fruit that's been lying around for years; Typesafe were the first to bother to stop and pick it up, and market the shit out of it. But it was still a low-hanging fruit.

      There are many, many of those left around, because AI frontier is moving forward so fast, everyone is racing ahead. Which is why I laugh when people say AI is not transformative and LLMs are a dead end (and my favorite, "what are we going to do with all those GPUs when the bubble pops?"). Even if SOTA LLMs hit a hard capability limit tomorrow and never advanced again, there's a good decade of growth and advancement to be extracted just from all the low-hanging fruits that were left unpicked along the way.

      1. seizethecheese · · focus · HN ↗
        Name a few of these low hanging fruit left around.
        1. TeMPOraL · · focus · HN ↗
          Jev is one.

          Diffusion transformers are not "easy" but underfunded.

          Random one in terms of applications: getting GPT-4-level[0] LLMs to operate at hundreds of tokens per second on edge hardware - opens up so many possibilities I'm probably unable to imagine half of them.

          E.g. Imagine spellcheck/predictive text (or code autocomplete) where the model is able to process a whole paragraph + surrounding application/system context in between keystrokes. Or an OS being able to reliably guess what you're doing in real-time, in between your UI interactions, and offer actually helpful contextual reactions.

          Or imagine finally funding some decent studies into exploring the models as computational artifacts - studying their latent spaces, how they form and how they model reality internally.

          Or imagine automated sliding doors that don't suck.

          --

          [0] - Or anything substantially better than BERT-level models used in Jev or that demo from the company doing inference ASICs, that has a chatbot online that does 14 kilotokens per second.

          1. aeve890 · · focus · HN ↗
            >Random one in terms of applications: getting GPT-4-level[0] LLMs to operate at hundreds of tokens per second on edge hardware

            That's low hanging for you?

            1. blurbleblurble · · focus · HN ↗
              It's likely quite close. There are so many papers proving concepts that would bring this, they just haven't been combined in production.
            2. msdz · · focus · HN ↗
              Maybe they meant in the sense of “untapped potential”, because so far a lot of the focus has been on increasing model capabilities, not necessarily performance/power budget.
              1. TeMPOraL · · focus · HN ↗
                Yes. Point is, it's untapped only because everyone is running in the race (even if out of curiosity), and there's just not enough people with means to tap into these side threads. For the past few years, there's been many interesting papers that circulated the industry, got recognized as worthwhile pursuits, and then dropped because running behind the Big Vendors had massively better ROI.
            3. ekabod · · focus · HN ↗
              That's a high hanging fruit, not low.
            4. guyomes · · focus · HN ↗
              If we throw in hardware dedicated to a specific LLM, it seems to be a rather low hanging fruit. Especially considering that this is already happening for vision models [1].

              [1]: &quot;FPGA-based CNN Acceleration using Pattern-Aware Pruning&quot; <a href="https:&#x2F;&#x2F;inria.hal.science&#x2F;hal-04689673&#x2F;document" rel="nofollow">https:&#x2F;&#x2F;inria.hal.science&#x2F;hal-04689673&#x2F;document

              1. mdp2021 · · focus · HN ↗
                &gt; hardware dedicated to a specific LLM

                That wording screams &quot;Taalas&quot;. Which, importantly, is not the only player trying to abate the distance between data and arithmetics...

            5. TeMPOraL · · focus · HN ↗
              Yes. It&#x27;s well within realm of possibility, but so far wasn&#x27;t pursued because the Big Vendors went all-in into capability growth (rightfully testing &quot;the bitter lesson&quot; to its limits) and got themselves stuck in an arms race, while everyone else is barely keeping up and&#x2F;or starstruck with fascination, exploring what these models can do.

              This got everyone racing forward and right now there is not enough human attention left in the world to productionize this, or any of the other &quot;side threads&quot;. When the race slows down, people will catch up, branch out, and loop back.

              1. blurbleblurble · · focus · HN ↗
                Just like renewable energy and so many other things. Hyperconcentration of capital is really tragic. I hope things turn around.
                1. TeMPOraL · · focus · HN ↗
                  They will. That&#x27;s the fallacy of the &quot;S-curve&quot; everyone likes to commit these days actually giving a positive outlook.

                  Assuming it won&#x27;t get to full RSI, the current approach will burn out - most likely economically. The race slows down, people branch out, look back, start picking up the &quot;untapped potential&quot;&#x2F;low-hanging fruits, and you have new S-curves launching in place of the one that just tapered off (hence a fallacy - a stack of S-curves adds up to continuing exponential growth).

                  In other words: it comes and goes. Hyperconcentrated capital will eventually deconcentrate.

            6. mdp2021 · · focus · HN ↗
              &gt; That&#x27;s low hanging for you

              An important part of the industry is studying that: it is built-up effort. Sooner or later, the fruits will be harvested. The targeted preparation has been there for years now.

            7. Twirrim · · focus · HN ↗
              We already have examples of LLMs running 16k+ tokens a second using custom ASICs.

              It&#x27;s down at the moment (Not sure if it&#x27;ll return?) but Chat Jimmy[0] produced by Taalas[1] was powered by an ASIC running Llama 3.1 8B, and hitting 17,000 tokens&#x2F;sec. It was amazing to use, you&#x27;d no sooner have hit enter than you had a full response back. I actually found its speed to be a problem for interactive stuff, as every answer got several paragraphs I&#x27;d then wade through, vs a model populating text closer to my reading speed.

              I appreciate, there are differences between an 8 billion parameter model and something GPT-4-ish, but we&#x27;re currently in the middle of a race between a half dozen or so companies to produce the next best frontier model, which requires their infrastructure to be dynamic.

              We really don&#x27;t always need newer better faster stronger models, there&#x27;s quite a lot of room for &quot;good enough&quot; where getting 17kt&#x2F;s at significantly lower power would be amazing.

              [0] <a href="https:&#x2F;&#x2F;chatjimmy.ai&#x2F;" rel="nofollow">https:&#x2F;&#x2F;chatjimmy.ai&#x2F; [1] <a href="https:&#x2F;&#x2F;taalas.com&#x2F;" rel="nofollow">https:&#x2F;&#x2F;taalas.com&#x2F;

              1. AshamedBadger56 · · focus · HN ↗
                &gt;I actually found its speed to be a problem for interactive stuff, as every answer got several paragraphs I&#x27;d then wade through, vs a model populating text closer to my reading speed.

                It would be interesting to pair the super fast model with a normal speed model. Have the super fast one do all the background research, code writing, etc. The normal model would just relay the needed info to you at a more reasonable pace.

                1. drewstiff · · focus · HN ↗
                  Why do you even need a second model for that? There are many simpler ways to slow down the rate that the near-immediate response is pseudo-typed onto the screen.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.