Clef: Open-weight decision models, and new RL fine-tuning platform
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Clef: Open-weight decision models, and new RL fine-tuning platform
Unofficial Hacker News client; not affiliated with Y Combinator.
manlymuppet · · focus · HN ↗
And it's only been a few weeks.
TeMPOraL · · focus · HN ↗
There are many, many of those left around, because AI frontier is moving forward so fast, everyone is racing ahead. Which is why I laugh when people say AI is not transformative and LLMs are a dead end (and my favorite, "what are we going to do with all those GPUs when the bubble pops?"). Even if SOTA LLMs hit a hard capability limit tomorrow and never advanced again, there's a good decade of growth and advancement to be extracted just from all the low-hanging fruits that were left unpicked along the way.
seizethecheese · · focus · HN ↗
murkt · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
TeMPOraL · · focus · HN ↗
Diffusion transformers are not "easy" but underfunded.
Random one in terms of applications: getting GPT-4-level[0] LLMs to operate at hundreds of tokens per second on edge hardware - opens up so many possibilities I'm probably unable to imagine half of them.
E.g. Imagine spellcheck/predictive text (or code autocomplete) where the model is able to process a whole paragraph + surrounding application/system context in between keystrokes. Or an OS being able to reliably guess what you're doing in real-time, in between your UI interactions, and offer actually helpful contextual reactions.
Or imagine finally funding some decent studies into exploring the models as computational artifacts - studying their latent spaces, how they form and how they model reality internally.
Or imagine automated sliding doors that don't suck.
--
[0] - Or anything substantially better than BERT-level models used in Jev or that demo from the company doing inference ASICs, that has a chatbot online that does 14 kilotokens per second.
aeve890 · · focus · HN ↗
That's low hanging for you?
blurbleblurble · · focus · HN ↗
msdz · · focus · HN ↗
TeMPOraL · · focus · HN ↗
ekabod · · focus · HN ↗
guyomes · · focus · HN ↗
[1]: "FPGA-based CNN Acceleration using Pattern-Aware Pruning" <a href="https://inria.hal.science/hal-04689673/document" rel="nofollow">https://inria.hal.science/hal-04689673/document
mdp2021 · · focus · HN ↗
That wording screams "Taalas". Which, importantly, is not the only player trying to abate the distance between data and arithmetics...
TeMPOraL · · focus · HN ↗
This got everyone racing forward and right now there is not enough human attention left in the world to productionize this, or any of the other "side threads". When the race slows down, people will catch up, branch out, and loop back.
blurbleblurble · · focus · HN ↗
TeMPOraL · · focus · HN ↗
Assuming it won't get to full RSI, the current approach will burn out - most likely economically. The race slows down, people branch out, look back, start picking up the "untapped potential"/low-hanging fruits, and you have new S-curves launching in place of the one that just tapered off (hence a fallacy - a stack of S-curves adds up to continuing exponential growth).
In other words: it comes and goes. Hyperconcentrated capital will eventually deconcentrate.
mdp2021 · · focus · HN ↗
An important part of the industry is studying that: it is built-up effort. Sooner or later, the fruits will be harvested. The targeted preparation has been there for years now.
Twirrim · · focus · HN ↗
It's down at the moment (Not sure if it'll return?) but Chat Jimmy[0] produced by Taalas[1] was powered by an ASIC running Llama 3.1 8B, and hitting 17,000 tokens/sec. It was amazing to use, you'd no sooner have hit enter than you had a full response back. I actually found its speed to be a problem for interactive stuff, as every answer got several paragraphs I'd then wade through, vs a model populating text closer to my reading speed.
I appreciate, there are differences between an 8 billion parameter model and something GPT-4-ish, but we're currently in the middle of a race between a half dozen or so companies to produce the next best frontier model, which requires their infrastructure to be dynamic.
We really don't always need newer better faster stronger models, there's quite a lot of room for "good enough" where getting 17kt/s at significantly lower power would be amazing.
[0] <a href="https://chatjimmy.ai/" rel="nofollow">https://chatjimmy.ai/ [1] <a href="https://taalas.com/" rel="nofollow">https://taalas.com/
AshamedBadger56 · · focus · HN ↗
It would be interesting to pair the super fast model with a normal speed model. Have the super fast one do all the background research, code writing, etc. The normal model would just relay the needed info to you at a more reasonable pace.
drewstiff · · focus · HN ↗
flipping_beacon · · focus · HN ↗
blurbleblurble · · focus · HN ↗
But also harnesses and more generally new insights on "the control flow problem" could end up squeezing a ton of performance out of small models.
Amekedl · · focus · HN ↗
Enough stuff can happen, software use itself might change, and that could really cause anything. "What will we do with all the gpus" might become a question if for a magnitude of tech and reasons leaked-opus-9 runs on a macbook m6 or 7
seizethecheese · · focus · HN ↗
dominotw · · focus · HN ↗
tomrod · · focus · HN ↗
alightsoul · · focus · HN ↗
[dead]
mdp2021 · · focus · HN ↗
Among the most important ones:
-- the long-known Problem of Transparency, applied to the apparent emergent intelligence in NNs. Why does it happen - in detail?
-- then, a Theory of Apparent Intelligence through NNs. Transforming the results achieved into a Science. Which allows to do what we are doing - but in a lean and targeted way.
-- then, a General Theory of Intelligence, that includes the above to go beyond current architectures and get those features of Intelligence we expect and still not have.
The long-term direction we got into must lead to this.
(You note a ponderant detail of the above when you note the importance of explaining the emergence of a World Model from a Language Model.)
TeMPOraL · · focus · HN ↗
patcon · · focus · HN ↗
hobofan · · focus · HN ↗
Of course you can also do ranking one-off with a decision model, but this likely less stable, and by doing pairwise ranking you can also relatively quickly do incremental inserts to the list.
sarkarghya · · focus · HN ↗
We have overcome split brain problems before so this wont be our first
esseph · · focus · HN ↗
CamperBob2 · · focus · HN ↗
There is still a lot we don't know about how to get the most out of existing LLM components from a speed or cognitive-performance perspective. People could easily spend the next decade studying and refining what's been built so far, even if no new, original approaches ever arrive.
randomNumber7 · · focus · HN ↗
sroussey · · focus · HN ↗
btown · · focus · HN ↗
Whether or not LLMs can self-improve their frontier capabilities, they can absolutely create a wake for themselves that accelerates everything else that's training on their synthetic data. We'll see every architecture of the past 40 years suddenly show leaps and bounds.
gutchapa · · focus · HN ↗
btown · · focus · HN ↗
What you can do with LLMs is reverse this: from any numbers of snapshots of flight data, you can create large numbers of plausible user queries, based on your data, that are known to be feasible or infeasible. And now you have a labeled data set to train a model that focuses solely on the query-creation and judgment systems. And you can experiment with whether having a more flexible query protocol leads to higher success rates without sacrificing accuracy, or whether you can generate that last-mile feasibility check as a combination of auditable code checks alongside AI-based judgment.
LLMs don't absolve you of having to break down your system architectures into components that have well-defined boundaries (though certainly they can help with that design). They do make those components feasible to solve at scale.
dzonga · · focus · HN ↗
then vision & robotics.
while everyone's chasing the frontier.
outofpaper · · focus · HN ↗
Good to see interest broadening beyond "just extend thinking." More approaches in the toolbox means fewer problems get treated as nails.
tomrod · · focus · HN ↗
Your comment here made me laugh, because I think we will finally be able to play Crysis at 10fps.
Just kidding, of course. GPU half lives are quite a bit less than standard compute half lives, no?[0] That's what I've been trying to understand regarding data centers focusing as GPU clusters -- seems like the ROI window would have to be very short for the capitalization.
[0] <a href="https://www.tomshardware.com/pc-components/gpus/datacenter-gpu-service-life-can-be-surprisingly-short-only-one-to-three-years-is-expected-according-to-unnamed-google-architect" rel="nofollow">https://www.tomshardware.com/pc-components/gpus/datacenter-g...