Clef: Open-weight decision models, and new RL fine-tuning platform
Thread
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
Clef: Open-weight decision models, and new RL fine-tuning platform
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
warkdarrior · · focus · HN ↗
kerenskiy · · focus · HN ↗
XTXinverseXTY · · focus · HN ↗
Moreover the specific prior art claim is absurd [1].
[0]: <a href="https://laya.convaiinnovations.com/" rel="nofollow">https://laya.convaiinnovations.com/
[1]: <a href="https://xtxinversexty.com/layas-prior-art-claim-is-absurd/" rel="nofollow">https://xtxinversexty.com/layas-prior-art-claim-is-absurd/
didibus · · focus · HN ↗
petercooper · · focus · HN ↗
theapadayo · · focus · HN ↗
The fascinating part to me is that Jev seems like this technique plus post-training to get multiple independent confidence values for each possible answer.
segmondy · · focus · HN ↗
woah · · focus · HN ↗
ford · · focus · HN ↗
zitterbewegung · · focus · HN ↗
ramoz · · focus · HN ↗
Anyone can copy that and apply to an array of models - stripped down LLMs or already slim/highly performant traditional classification architectures (just wrap inference with an api that inputs/outputs the same structured data).
Jev, I think, would say their advantage is the intelligence of their models and training data including calibration: <a href="https://medium.com/code-applied/calibrated-classifiers-making-your-models-probabilities-trustworthy-8ccbd07c86a7" rel="nofollow">https://medium.com/code-applied/calibrated-classifiers-makin... (which i still struggle with in the general application... there's no free lunch with these things).
233mhz · · focus · HN ↗
If you have a very narrow use case you can train a BERT based decision model on a laptop an hour if you have good data to train it on. It'll answer faster than the roundtrip to clef/jev and use <1gb memory
conmod278 · · focus · HN ↗
porridgeraisin · · focus · HN ↗
Getting training data that works well for calibrated classification objectives is difficult.
I hear conflicting opinions (including my own) about how well calibrated each of these are. Jev seems to be the best.
But the jev release made obvious the PMF for these models, and the underlying reality is that calibration really doesn't matter much when you're replacing usecases where people were using damn LM head softmax probabilities before, which are nowhere near calibrated.
So now everyone simply finetunes qwen and makes a compared-to-regular-LLM vastly cheaper decision model. And it works for majority of usecases.
pizzafeelsright · · focus · HN ↗
Many people seem to have run into the same question and started working out the answer.
giancarlostoro · · focus · HN ↗
It seems insanely obvious at least to me, that JEV is the new hot thing for the AI field since they give you stronger output that isn't... flat out wrong, that alone is impressive.
nico · · focus · HN ↗
But, for these adhoc models, you need to understand the task more, collect some data and train the model (on CPU, no need for GPU). So Jev-like models are a great way of getting a hosted general decision model, but if you have a very narrow task or set of tasks, you might be better off with some more basic models that you can run on the same server you run other things or even on your laptop
johnecheck · · focus · HN ↗
hbcdbff · · focus · HN ↗
alashow · · focus · HN ↗
thih9 · · focus · HN ↗
Isn’t this a potential attack vector?
aryabakh · · focus · HN ↗
ssiddharth · · focus · HN ↗
CBLT · · focus · HN ↗
DesaiAshu · · focus · HN ↗
open592 · · focus · HN ↗
swingboy · · focus · HN ↗
ttul · · focus · HN ↗
yipinwong · · focus · HN ↗
Or are companies/people already building this based on say an arXiv docs? n
---
The pricing is ... hm more expensive but not at the point I won't give it a try due to the embeded vision encoding
XCSme · · focus · HN ↗
Latency won't be that good, but could still work similarly. Simply force the structured output of a LLM to the given schema.
Probably also easy to train because we can use stronget LLMs to generate input/output data, or even synthetic data is easy to generate.
It's not really a new technology, it's more like a new use-case.
sigbottle · · focus · HN ↗
popinman322 · · focus · HN ↗
<a href="https://github.com/blockbrain-ai/cygnet-recipe" rel="nofollow">https://github.com/blockbrain-ai/cygnet-recipe
orbital-decay · · focus · HN ↗
redox99 · · focus · HN ↗
redox99 · · focus · HN ↗
conmod278 · · focus · HN ↗
<a href="https://www.youtube.com/watch?v=AzxoU7kxjig" rel="nofollow">https://www.youtube.com/watch?v=AzxoU7kxjig
nico · · focus · HN ↗
I've been playing with this for the last year or so. Started with a personal email classifier, also did benchmarks with some public datasets, then created a couple classifiers that could play Doom, and now I've been trying out some other experiments, like a request proxy/router to automatically choose a classifier and fallback to LLM to handle unseen requests
Jev did a great job at creating hype, but also at shaping the concept and space of "decision engine" or "decision model". People were already doing this with LLMs, which is very inefficient for most tasks like that, and the Jev guys figured there was a market there. It seems like they were right, and now there's a rush to flood the space, taking advantage of the hype window
calebkaiser · · focus · HN ↗
In general, training a general purpose classifier is something lots of people have worked on for a long time. Large Transformer models themselves are typically "generalists" already, so structured generation and constrained decoding have given you the ability to use an LLM as a general classifier for years. It's an incredibly common pattern for working with LLM judges or any sort of branched decision making workflow.
A lot of people who are a bit less familiar with the field saw the hype around Jev and presumed that the reason it was so exciting was that it was a fundamentally new interface for working with an LLM. And that additional excitement drove even more attention to Jev. But fundamentally, TypeSafe's announcement was that they found a particular architecture/training paradigm that resulted in a model for this particular interface that had incredible accuracy, very low latency, and for which they could offer inference at a super low cost.
I've not kept up with the flood of Jev clones that have been released, but I think this is just typical for any new component in deep learning that gets popular. There are an absurd number of open source autoregressive LLMs and fine tunes you can use. The thing that makes one more popular than the other is typically the general performance of the individual model.
But training a model for this purpose, or emulating the procedures described in Jev's papers, isn't something that would be beyond the capabilities of any lab. It's not an entirely alien architecture or approach.
The bigger question for TypeSafe as a company would be if other teams are producing Jev-like models that win on performance or cost. Like I said, I haven't followed the reports super closely, so no idea if that's the case or not.
tomrod · · focus · HN ↗
If CF's benchmark is representative and sufficient, Clef outperforms Jev!
Models by themselves don't guarantee market capture. Rather, its how they integrate. I think a lot of folks are burned by the closed nature of many models.
TeMPOraL · · focus · HN ↗
Now that we're hitting against the hardware supply limits of global economy, I expect more people to go back and revisit the things left along the way in the mad rush to "just throw more compute at it / make a bigger model" - and thus many more cases like Jev to show up in the next few years.
orbital-decay · · focus · HN ↗
janalsncm · · focus · HN ↗
The hard part is the data and evaluation. Sure, it’s not that hard to build a fast model with good predictive power. But fast at doing what? You probably don’t care about classifying whether a hotdog is a sandwich (which is the Jev demo).
zwaps · · focus · HN ↗
kflansburg · · focus · HN ↗
bityard · · focus · HN ↗
NitpickLawyer · · focus · HN ↗
There is no official qwen 3.8 9b
ddarolfi · · focus · HN ↗
verdverm · · focus · HN ↗
bityard · · focus · HN ↗
okpatil · · focus · HN ↗
16ms latency. And locally run.
<a href="https://at0m.pienomial.com/" rel="nofollow">https://at0m.pienomial.com/
Why go big when you can go small ?
kamranjon · · focus · HN ↗
okpatil · · focus · HN ↗
To counter, most of the AI is not open. So is none of Microsoft Products. As long as they work, we keep using them.
verdverm · · focus · HN ↗
Ai is too important and transformational to let Big Ai dominate in a closed ecosystem
okpatil · · focus · HN ↗
ricardobeat · · focus · HN ↗
okpatil · · focus · HN ↗
MisterMunchkin · · focus · HN ↗
RGS1811 · · focus · HN ↗
hansonkd · · focus · HN ↗
Instead, to make up for the lack of economic viability of their models, they are forced to release publicly to get marketing to get others to pay based on hype.
globular-toast · · focus · HN ↗
6thbit · · focus · HN ↗
Perhaps that may be too costly atm
manlymuppet · · focus · HN ↗
And it's only been a few weeks.
TeMPOraL · · focus · HN ↗
There are many, many of those left around, because AI frontier is moving forward so fast, everyone is racing ahead. Which is why I laugh when people say AI is not transformative and LLMs are a dead end (and my favorite, "what are we going to do with all those GPUs when the bubble pops?"). Even if SOTA LLMs hit a hard capability limit tomorrow and never advanced again, there's a good decade of growth and advancement to be extracted just from all the low-hanging fruits that were left unpicked along the way.
seizethecheese · · focus · HN ↗
murkt · · focus · HN ↗
MadrasTh0rn · · focus · HN ↗
TeMPOraL · · focus · HN ↗
Diffusion transformers are not "easy" but underfunded.
Random one in terms of applications: getting GPT-4-level[0] LLMs to operate at hundreds of tokens per second on edge hardware - opens up so many possibilities I'm probably unable to imagine half of them.
E.g. Imagine spellcheck/predictive text (or code autocomplete) where the model is able to process a whole paragraph + surrounding application/system context in between keystrokes. Or an OS being able to reliably guess what you're doing in real-time, in between your UI interactions, and offer actually helpful contextual reactions.
Or imagine finally funding some decent studies into exploring the models as computational artifacts - studying their latent spaces, how they form and how they model reality internally.
Or imagine automated sliding doors that don't suck.
--
[0] - Or anything substantially better than BERT-level models used in Jev or that demo from the company doing inference ASICs, that has a chatbot online that does 14 kilotokens per second.
aeve890 · · focus · HN ↗
That's low hanging for you?
blurbleblurble · · focus · HN ↗
msdz · · focus · HN ↗
TeMPOraL · · focus · HN ↗
ekabod · · focus · HN ↗
guyomes · · focus · HN ↗
[1]: "FPGA-based CNN Acceleration using Pattern-Aware Pruning" <a href="https://inria.hal.science/hal-04689673/document" rel="nofollow">https://inria.hal.science/hal-04689673/document
mdp2021 · · focus · HN ↗
That wording screams "Taalas". Which, importantly, is not the only player trying to abate the distance between data and arithmetics...
TeMPOraL · · focus · HN ↗
This got everyone racing forward and right now there is not enough human attention left in the world to productionize this, or any of the other "side threads". When the race slows down, people will catch up, branch out, and loop back.
blurbleblurble · · focus · HN ↗
TeMPOraL · · focus · HN ↗
Assuming it won't get to full RSI, the current approach will burn out - most likely economically. The race slows down, people branch out, loop back, pick up the "untapped potential"/low-hanging fruits, and you have new S-curves launching in place of the one that just tapered off (hence a fallacy - a stack of S-curves adds up to continuing exponential growth).
mdp2021 · · focus · HN ↗
An important part of the industry is studying that: it is built-up effort. Sooner or later, the fruits will be harvested. The targeted cultivation has been there for years now.
Twirrim · · focus · HN ↗
It's down at the moment (Not sure if it'll return?) but Chat Jimmy[0] produced by Taalas[1] was powered by an ASIC running Llama 3.1 8B, and hitting 17,000 tokens/sec. It was amazing to use, you'd no sooner have hit enter than you had a full response back. I actually found its speed to be a problem for interactive stuff, as every answer got several paragraphs I'd then wade through, vs a model populating text closer to my reading speed.
I appreciate, there are differences between an 8 billion parameter model and something GPT-4-ish, but we're currently in the middle of a race between a half dozen or so companies to produce the next best frontier model, which requires their infrastructure to be dynamic.
We really don't always need newer better faster stronger models, there's quite a lot of room for "good enough" where getting 17kt/s at significantly lower power would be amazing.
[0] <a href="https://chatjimmy.ai/" rel="nofollow">https://chatjimmy.ai/ [1] <a href="https://taalas.com/" rel="nofollow">https://taalas.com/
AshamedBadger56 · · focus · HN ↗
It would be interesting to pair the super fast model with a normal speed model. Have the super fast one do all the background research, code writing, etc. The normal model would just relay the needed info to you at a more reasonable pace.
drewstiff · · focus · HN ↗
flipping_beacon · · focus · HN ↗
blurbleblurble · · focus · HN ↗
Amekedl · · focus · HN ↗
Enough stuff can happen, software use itself might change, and that could really cause anything. "What will we do with all the gpus" might become a question if for a magnitude of tech and reasons leaked-opus-9 runs on a macbook m6 or 7
seizethecheese · · focus · HN ↗
dominotw · · focus · HN ↗
tomrod · · focus · HN ↗
alightsoul · · focus · HN ↗
[dead]
mdp2021 · · focus · HN ↗
Among the most important ones:
-- the long-known Problem of Transparency, applied to the apparent emergent intelligence in NNs. Why does it happen - in detail?
-- then, a Theory of Apparent Intelligence through NNs. Transforming the results achieved into a Science. Which allows to do what we are doing - but in a lean and targeted way.
-- then, a General Theory of Intelligence, that includes the above to go beyond current architectures and get those features of Intelligence we expect and still not have.
The long-term direction we got into must lead to this.
(You note a ponderant detail of the above when you note the importance of explaining the emergence of a World Model from a Language Model.)
TeMPOraL · · focus · HN ↗
patcon · · focus · HN ↗
hobofan · · focus · HN ↗
Of course you can also do ranking one-off with a decision model, but this likely less stable, and by doing pairwise ranking you can also relatively quickly do incremental inserts to the list.
sarkarghya · · focus · HN ↗
We have overcome split brain problems before so this wont be our first
esseph · · focus · HN ↗
CamperBob2 · · focus · HN ↗
There is still a lot we don't know about how to get the most out of existing LLM components from a speed or cognitive-performance perspective. People could easily spend the next decade studying and refining what's been built so far, even if progress halted tomorrow.
randomNumber7 · · focus · HN ↗
sroussey · · focus · HN ↗
btown · · focus · HN ↗
Whether or not LLMs can self-improve their frontier capabilities, they can absolutely create a wake for themselves that accelerates everything else that's training on their synthetic data. We'll see every architecture of the past 40 years suddenly show leaps and bounds.
gutchapa · · focus · HN ↗
btown · · focus · HN ↗
What you can do with LLMs is reverse this: from any numbers of snapshots of flight data, you can create large numbers of plausible user queries, based on your data, that are known to be feasible or infeasible. And now you have a labeled data set to train a model that focuses solely on the query-creation and judgment systems. And you can experiment with whether having a more flexible query protocol leads to higher success rates without sacrificing accuracy, or whether you can generate that last-mile feasibility check as a combination of auditable code checks alongside AI-based judgment.
LLMs don't absolve you of having to break down your system architectures into components that have well-defined boundaries (though certainly they can help with that design). They do make those components feasible to solve at scale.
dzonga · · focus · HN ↗
then vision & robotics.
while everyone's chasing the frontier.
outofpaper · · focus · HN ↗
Good to see interest broadening beyond "just extend thinking." More approaches in the toolbox means fewer problems get treated as nails.
tomrod · · focus · HN ↗
Your comment here made me laugh, because I think we will finally be able to play Crysis at 10fps.
Just kidding, of course. GPU half lives are quite a bit less than standard compute half lives, no?[0] That's what I've been trying to understand regarding data centers focusing as GPU clusters -- seems like the ROI window would have to be very short for the capitalization.
[0] <a href="https://www.tomshardware.com/pc-components/gpus/datacenter-gpu-service-life-can-be-surprisingly-short-only-one-to-three-years-is-expected-according-to-unnamed-google-architect" rel="nofollow">https://www.tomshardware.com/pc-components/gpus/datacenter-g...
slopnt · · focus · HN ↗
alightsoul · · focus · HN ↗
smallmancontrov · · focus · HN ↗
"Discriminative" always had pointlessly bad optics, but I knew it was over when I started seeing prominent machine learning researchers who p=100% knew better describe discriminative models as generative because that was the buzzword of the year. "Decision model" sells the value proposition much better and doesn't sound like an anti-woke crusade.
tomrod · · focus · HN ↗
eastdakota · · focus · HN ↗
segmondy · · focus · HN ↗
SebastianSosa · · focus · HN ↗
Foobar8568 · · focus · HN ↗
verdverm · · focus · HN ↗
- I'd like a option, are you sure, not, repeat
- in and out of doors
- sisyphean effort in the cave
- jev-ish level grinding
It was impressive, beat pomemon for less than $2, but not all that interesting. People asking how different Math.Random plays pokemon would be
indoor47 · · focus · HN ↗
"Clef builds upon this concept, but uses a different base model as the backbone. We currently use Qwen as the base model and post-trained it to suit decision model use cases. "
heliosAtwork · · focus · HN ↗
"In the same week that Jev came out, we posted about some experiments [1] we had with our own homegrown decision model."
[1] <a href="https://x.com/michellechen/status/2101091012559151480" rel="nofollow">https://x.com/michellechen/status/2101091012559151480
lofaszvanitt · · focus · HN ↗
fwip · · focus · HN ↗
As far as I've seen, all of the Jev-compatible projects simply take an LLM, hack off a layer or two at the end, and call it good. Some of them spend more work than others trying to back-estimate in accurate probabilities.
rahimnathwani · · focus · HN ↗
A) Structured outputs from LLMs (doesn't need fine tuning but can be expensive)
B) Classification output from fine-tuned BERT-like or GLiNER models (is calibrated well and has cheap/fast inference)
What Jev did is combine the advantages of both A and B into one model/product, and create a really good API.
They claim that a key innovation is how they've trained the model using what they call RLCD (RL from calibrated decisions). So it's not just that you can get the outputs (which is easy to add to any LLM) but that the different primitives they expose (Choice, Score, Noul) have each been calibrated. For example, they claim that if you use the Score primitive (which gives you probabilities along a bunch of choices representing a continuum) that's not just using the more general 'Choice' primitive under the hood. It's been calibrated separately.
I don't know how many of the Jev-like things we've seen do that. For example Cloudflare offers a Jev-like model with the same API, and which they say was trained with RLCD. But I don't know whether Choice and Score are different under the hood, or whether Score is just sugar on top of Choice. (Should be easy to test this, but I haven't done it.)
nater5000 · · focus · HN ↗
You make it sound like this is some noteworthy timeframe. I'd be more surprised if Jev wasn't immediately made obsolete within days of release (which was effectively the case), especially when backed by a company like Cloudflare lol
ML moves quick. Add the extreme hype and cash floating around in this space and you can expect that anything resembling something novel and relatively untapped is going to be pounced on and turned over basically immediately.
xnx · · focus · HN ↗
manlymuppet · · focus · HN ↗
With Jev, structured output is its native format. Jev is just way faster, cheaper, and outright better for a lot of things.
Note that while all of this is great, this is nowhere near a sort of "ChatGPT moment". It's a cool new thing, and it's way better at certain tasks, which could be big. That's all though.
buildbuildbuild · · focus · HN ↗
The weights have permissive licensing, but the data and training pipeline are not published to reproduce them from their proprietary Qwen starting points. Weights are not "source."
jMyles · · focus · HN ↗
<sad trombone sound>
Surely someone will soon do what the title of this post makes it seem like cloudfare did. Truly modular open source training and inference logic, along with a totally open corpus and weights, will eventually out-compete the closed ecosystem.
ainch · · focus · HN ↗
dang · · focus · HN ↗
ksymph · · focus · HN ↗
reassess_blind · · focus · HN ↗
schainks · · focus · HN ↗
esafak · · focus · HN ↗
selfawareMammal · · focus · HN ↗
mococa · · focus · HN ↗
swe_dima · · focus · HN ↗
jasfi · · focus · HN ↗
I built this for my own needs, and thought others might find it useful too.
croemer · · focus · HN ↗
handfuloflight · · focus · HN ↗
croemer · · focus · HN ↗
mpolichette · · focus · HN ↗
I'd love an privacy first on-device model i could use in iOS.
okpatil · · focus · HN ↗
At0M: A 60M local Jev at 16 ms latency and 79% accuracy on Typed Decision
<a href="https://at0m.pienomial.com/" rel="nofollow">https://at0m.pienomial.com/ <a href="https://news.ycombinator.com/item?id=49920350">https://news.ycombinator.com/item?id=49920350
itsmeduncan · · focus · HN ↗
[dead]
okpatil · · focus · HN ↗
At0M: A 60M local Jev at 16 ms latency and 79% accuracy on Typed Decision
<a href="https://at0m.pienomial.com/" rel="nofollow">https://at0m.pienomial.com/ <a href="https://news.ycombinator.com/item?id=49920350">https://news.ycombinator.com/item?id=49920350
networked · · focus · HN ↗
You can list the decisions models available on OpenRouter: <a href="https://openrouter.ai/models?output_modalities=decisions" rel="nofollow">https://openrouter.ai/models?output_modalities=decisions. There are currently nine.
A quick benchmark for my AI harness (<a href="https://github.com/dbohdan/strument/tree/c2320ca142f77495ecdcadd3cd7a4c9b9445369f/doc/experiments/2026-10-decision-models" rel="nofollow">https://github.com/dbohdan/strument/tree/c2320ca142f77495ecd...) showed that of five models on OpenRouter only Liquid D1 performed similar to Jev for determining command safety.
damsta · · focus · HN ↗
fooker · · focus · HN ↗
I bet the competition will result in research into how to make these decision models several more orders of magnitude faster and cheaper.
Here's a challenge problem - look at a 1M context window and produce N decisions (different queries) from it in 50-100ms.
mrkn1 · · focus · HN ↗
fastball · · focus · HN ↗
The value isn't really in the I/O shape, it is in the intelligence combined with the output shape. Every extra ounce of intelligence in these models unlocks additional use-cases. But the converse is also true: a dumb decision model is going to be less useful than using a more intelligent standard LLM.
That is the appeal of Jev: for certain usage it has more intelligence than some small SOTA LLMs. It is the first decision model that actually feels intelligent (to me).
mrkn1 · · focus · HN ↗
ralusek · · focus · HN ↗
Jev/TypeSafe: 230 ms median, 254 ms mean Jev/OpenRouter: 237 ms median, 267 ms mean Clef Flash: 661 ms median, 806 ms mean
What gives?
alex7o · · focus · HN ↗
okpatil · · focus · HN ↗
At0M: A 60M local Jev at 16 ms latency and 79% accuracy on Typed Decision
<a href="https://at0m.pienomial.com/" rel="nofollow">https://at0m.pienomial.com/ <a href="https://news.ycombinator.com/item?id=49920350">https://news.ycombinator.com/item?id=49920350
afzalive · · focus · HN ↗
okpatil · · focus · HN ↗
Businesses are built on outliers. It doesn't make sense throwing your hard earned insights while paying them money to steal it.
Also, cloudflare <a href="https://robindev.substack.com/p/cloudflare-took-down-our-website" rel="nofollow">https://robindev.substack.com/p/cloudflare-took-down-our-web...
dcastm · · focus · HN ↗
okpatil · · focus · HN ↗
cootsnuck · · focus · HN ↗
okpatil · · focus · HN ↗
We wanted to stress test the system before the V1 release.
Ciph · · focus · HN ↗
okpatil · · focus · HN ↗
kamranjon · · focus · HN ↗
okpatil · · focus · HN ↗
Apologies if it is too much of a bother.
Transformanshen · · focus · HN ↗
okpatil · · focus · HN ↗
curl -s -X POST <a href="https://at0m.pienomial.com/decide/v0" rel="nofollow">https://at0m.pienomial.com/decide/v0 \ -H 'Content-Type: application/json' \ -d '{ "state": "Charged twice for the same card payment this morning.", "questions": { "queue": {"type":"choice", "instructions":"Which team should handle this?", "criteria": {"billing":"invoices, charges, refunds", "technical":"outages, bugs, deploys", "fraud":"unauthorised or suspicious activity"}}, "urgent": {"type":"noul", "instructions":"Needs action today."}}, "email_id": "you@example.com"}'
If it fits your use case, you are welcome to use it.
When it is a rust standalone rust executable, as it is powering the API, it becomes just plug and play. No dependencies needed.
ranyume · · focus · HN ↗
amluto · · focus · HN ↗
I’d love to see someone build a model of this sort that can actually accept priors and do something intelligent with them.
[0] You can feed Jev a prior as text. I’ve tried it. It works poorly.
brokensegue · · focus · HN ↗
okpatil · · focus · HN ↗
We believe entire compliance workflows (even multilingual) could be automated.
Would you like to get a demo ?
sheepscreek · · focus · HN ↗
okpatil · · focus · HN ↗
It is possible with deterministic decision models, such as At0m, to gauge the probabilities at every decision. This behavior in addition to hard coded logic, it is possible to completely replicate a prompt's logic.
Using Fable 5.1, it is a matter of minutes.
I believe that most of the compliance check documents will be a solved problem, 3-6 months in future.
None of the LLMs can do it.
Hence I asked to the comment poster if he would want to demo, so that I can show it to him, how to do it step by step. By bad, if it came out too strongly.
sheepscreek · · focus · HN ↗
amluto · · focus · HN ↗
> Confidence is derived from the probabilities
<a href="https://docs.typesafe.ai/confidence" rel="nofollow">https://docs.typesafe.ai/confidence
My inner Bayesian would like for Jev to provide something resembling “evidence”, although I admit that one might ask Jev questions that are somewhat awkward to treat as typical Bayesian questions. If I ask “will this PR be merged”, it’s kind of strange to contemplate the probability of a PR conditioned in that PR being merged in the future. But I bet there is a way to formalize a prior-free classifier in a way that makes Bayesians and non-Bayesians happy, possibly involving actual learned probabilities and confidence levels. If you read the literature on scoring rules, you will find that classifier scores do somewhat naturally decompose into a few interpretable terms.
mikeocool · · focus · HN ↗
If I have to gather and tag data to fine-tune Jev, I can probably just train an "old school" classifier model and make it even cheaper, faster, and just as accurate.
jdthedisciple · · focus · HN ↗
amluto · · focus · HN ↗
AnthusAI · · focus · HN ↗
You can also improve your Jev classifications based on your ongoing data if you're labeling it continuously, especially if you're explaining the reasoning in the feedback labels. You can identify new elements of the rubric and add them to the list of classifications that Jev produces, and then those become new features for your ML model.
Two levers of control for using data to make a Jev-based classifier model continuously better-aligned.
AnthusAI · · focus · HN ↗
wakeywakeywakey · · focus · HN ↗
nikcub · · focus · HN ↗
gitghxst · · focus · HN ↗
cakoose · · focus · HN ↗
1. Humans are already not in the loop for lots of LLM agent actions. Isn't that just a function of how much you trust it and not some completely new paradigm? Am I missing something?
2. How can it gather context if it just outputs a single decision?
One guess: Maybe it's decision can be "gather more context and re-run me"? But an LLM can be much more expressive about what context it needs.
lmc · · focus · HN ↗
vulture916 · · focus · HN ↗
At 300 tokens per call, you'd get:
One million decisions on Jev cost about $12.60. One million decisions on Clef cost about $72.
Would probably make sense to self-host Clef, if you have the capability/resources. If not...
jampekka · · focus · HN ↗
verdverm · · focus · HN ↗
scronkfinkle · · focus · HN ↗
It's weird to think of these kinds of models as having "output tokens". Cross-encoder approaches like Laya add a [MASK] marker per option, but nothing is generated the way an autoregressive transformer generates. It's one bidirectional pass over your input, then a small head scores each option, so you wouldn't really pay for output as much as only input
strangescript · · focus · HN ↗
Since it benches better than Jev, Jev is probably smaller and easier to host. They could also be losing a lot of money.
wongarsu · · focus · HN ↗
And while Jev has a lot of latency that latency stays pretty flat with larger inputs, so it's probably related to their inference pipeline, not model size
andrewingram · · focus · HN ↗
Havoc · · focus · HN ↗
For privacy perhaps but on pricing you're unlikely to come out ahead versus datacentres with scale and industrial power pricing
otabdeveloper4 · · focus · HN ↗
Turns out AWS is actually about 10 times more expensive than renting your own datacenter compute.
Havoc · · focus · HN ↗
It’s like comparing pricing of a ribeye at the butcher and on restaurant menu. It’s apples and orange.
Certainly AWS is making good money too though. And yeah if you just need a VM then AWS isn’t the way
winddude · · focus · HN ↗
meander_water · · focus · HN ↗
This seems misleading. Decision models do not produce deterministic output. Repeated calls can product different decisions just like an LLM with structured outputs.
lmc · · focus · HN ↗
_superposition_ · · focus · HN ↗
adobe · · focus · HN ↗
[dead]
adobe · · focus · HN ↗
[dead]
KieranGrosvenor · · focus · HN ↗
[dead]
radium3d · · focus · HN ↗
pdlug · · focus · HN ↗
Quality: close (recall 0.98 vs 1.00) Hosted p50: Clef ~850ms, Jev ~110ms Clef-flash: over-escalates
Data + script: <a href="https://github.com/nicia-ai/admission-decision-eval" rel="nofollow">https://github.com/nicia-ai/admission-decision-eval
Wazzymandias · · focus · HN ↗
This entire time they could have just said "decision model" but they kept using vague, flowery wording. I have no idea why
agrippanux · · focus · HN ↗
Clef was 2-3x slower and worse (it caught less hate speech) than Jev. Overall disappointing.
teleforce · · focus · HN ↗
Although they included the benchmark against DJev and Clef is better, perhaps if you can test it to see the real-world performance.
nostrebored · · focus · HN ↗
wongarsu · · focus · HN ↗
For small inputs Jev is slow, but its latency curve is very flat. Fine-tuning a decently-sized llm (like this 27B model) gives you something that's faster on small input sizes, but even with moderate contexts quickly becomes much slower than Jev. Characterizing it as "faster than Jev" is very misleading, unless you know all your questions have tiny context (less than 1k tokens or so)
tomrod · · focus · HN ↗
wongarsu · · focus · HN ↗
reexpressionist · · focus · HN ↗
[dead]
lin7c · · focus · HN ↗
[dead]
johnbatch · · focus · HN ↗
AmazingTurtle · · focus · HN ↗
Maxforever · · focus · HN ↗
[dead]
ricardobeat · · focus · HN ↗
tough · · focus · HN ↗
Jeeetendra · · focus · HN ↗
pcthrowaway · · focus · HN ↗
Wait, does this mean this is the first (non-generative) model to be able to do <a href="https://xkcd.com/1425/" rel="nofollow">https://xkcd.com/1425/ !?
bullcitydev · · focus · HN ↗
Gshaheen · · focus · HN ↗
dgacmu · · focus · HN ↗
leopoldj · · focus · HN ↗
1. <a href="https://huggingface.co/solanaclawd/clef-solana-research" rel="nofollow">https://huggingface.co/solanaclawd/clef-solana-research
ActivePattern · · focus · HN ↗
If it needs to be fine-tuned, then why not fine-tune smaller and cheaper models that are only as big as the task requires?
leopoldj · · focus · HN ↗
dgacmu · · focus · HN ↗
djray · · focus · HN ↗
verdverm · · focus · HN ↗
<a href="https://blog.cloudflare.com/ai-code-review/" rel="nofollow">https://blog.cloudflare.com/ai-code-review/
sauercrowd · · focus · HN ↗
iugtmkbdfil834 · · focus · HN ↗
ramoz · · focus · HN ↗
iugtmkbdfil834 · · focus · HN ↗
Sorry, sometimes I forget not everyone is me;p
Hmm,where do you think jev could fit?
pwiklacz · · focus · HN ↗
[dead]
itsmeduncan · · focus · HN ↗
[dead]