Clef: Open-weight decision models, and new RL fine-tuning platform
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Clef: Open-weight decision models, and new RL fine-tuning platform
Unofficial Hacker News client; not affiliated with Y Combinator.
manlymuppet · · focus · HN ↗
And it's only been a few weeks.
TeMPOraL · · focus · HN ↗
There are many, many of those left around, because AI frontier is moving forward so fast, everyone is racing ahead. Which is why I laugh when people say AI is not transformative and LLMs are a dead end (and my favorite, "what are we going to do with all those GPUs when the bubble pops?"). Even if SOTA LLMs hit a hard capability limit tomorrow and never advanced again, there's a good decade of growth and advancement to be extracted just from all the low-hanging fruits that were left unpicked along the way.
seizethecheese · · focus · HN ↗
TeMPOraL · · focus · HN ↗
Diffusion transformers are not "easy" but underfunded.
Random one in terms of applications: getting GPT-4-level[0] LLMs to operate at hundreds of tokens per second on edge hardware - opens up so many possibilities I'm probably unable to imagine half of them.
E.g. Imagine spellcheck/predictive text (or code autocomplete) where the model is able to process a whole paragraph + surrounding application/system context in between keystrokes. Or an OS being able to reliably guess what you're doing in real-time, in between your UI interactions, and offer actually helpful contextual reactions.
Or imagine finally funding some decent studies into exploring the models as computational artifacts - studying their latent spaces, how they form and how they model reality internally.
Or imagine automated sliding doors that don't suck.
--
[0] - Or anything substantially better than BERT-level models used in Jev or that demo from the company doing inference ASICs, that has a chatbot online that does 14 kilotokens per second.
mdp2021 · · focus · HN ↗
Among the most important ones:
-- the long-known Problem of Transparency, applied to the apparent emergent intelligence in NNs. Why does it happen - in detail?
-- then, a Theory of Apparent Intelligence through NNs. Transforming the results achieved into a Science. Which allows to do what we are doing - but in a lean and targeted way.
-- then, a General Theory of Intelligence, that includes the above to go beyond current architectures and get those features of Intelligence we expect and still not have.
The long-term direction we got into must lead to this.
(You note a ponderant detail of the above when you note the importance of explaining the emergence of a World Model from a Language Model.)
TeMPOraL · · focus · HN ↗
patcon · · focus · HN ↗