It can't hallucinate, but it doesn't mean it can't make wrong decisions. Just because it adheres to a specific output format at all time, while LLMs have the output format at their mercy, then the claim of not hallucinating is made technically true.
I think that this specific part is not super interesting if your harness just recovers from invalid LLM outputs.
The latency and cost - yes, those are super interesting.
You can get rigid output format from "classic" LLMs <a href="https://docs.vllm.ai/en/latest/features/structured_outputs/" rel="nofollow">https://docs.vllm.ai/en/latest/features/structured_outputs/ though model support is limited.
Would like to have something like in the original post but open weights.
The model can't reason comprehensively (e.g., like Sol XHigh would to solve a complicated problem), but it's designed to be able to answer anything a human reasonably could quickly and intuitively, i.e., system one thinking: <a href="https://en.wikipedia.org/wiki/Thinking,_Fast_and_Slow" rel="nofollow">https://en.wikipedia.org/wiki/Thinking,_Fast_and_Slow
I was wondering the same! So I asked Astra to build me <a href="https://jev-chess-master.vercel.app/" rel="nofollow">https://jev-chess-master.vercel.app/ where you play against Jev AI as chess player.
You play White, and Jev plays Black. Rather than asking an LLM to generate move strings or JSON, the backend feeds all server-validated legal candidate moves into Vercel AI SDK's experimental_evaluate(). Jev picks Black's move and outputs its probability distribution across all legal candidates in a single forward pass (~300ms, ~$0.00004/move).
I am horrible at chess and easily won. It does play very very badly.
But it is essentially playing bullet chess right? No time to reason and is basically forced to go with it's knee jerk reaction. So maybe it plays reasonably against another average bullet chess player?
This is an instructive demo and makes me pause a bit when thinking about where I would trust using a classifier like this. Also there is no way to fine tune it right? So you are just left to its interpretation of the classification schema.
Building classifiers is hard and forces you into thinking about your problem space and your comfort level with type1 or type2 errors. I fear this will encourage sloppy work because all those decisions are black boxed.
Not like we were in a utopia careful ML applications before this.
I had a different initial confusion - it seems this company has no relation to the company formerly known as Typesafe <a href="https://en.wikipedia.org/wiki/Akka.io" rel="nofollow">https://en.wikipedia.org/wiki/Akka.io
I definitely agree it's underexplained in type safe.ai's materials.
I have to assume it's a reference to the fast, heuristic, intuitive "system 1" process in humans, as opposed to the slow, procedural, reasoning "system 2".
This theory is recognized, among others, in Daniel Kahneman 2002 Nobel prize on Economics.
Struggled to understand the use case till after i came across this <a href="https://eu.36kr.com/en/p/3991165197188099" rel="nofollow">https://eu.36kr.com/en/p/3991165197188099 article, which referred to them as the "referee for agents."
dinobones · · focus · HN ↗
Typesafe.AI sounds like some typescript/structured output type of tool…
What even is “system one” ?
IMO the product/tech is really there, just needs better communication.
toddmorey · · focus · HN ↗
flyinglizard · · focus · HN ↗
I think that this specific part is not super interesting if your harness just recovers from invalid LLM outputs.
The latency and cost - yes, those are super interesting.
mhitza · · focus · HN ↗
Would like to have something like in the original post but open weights.
zenlikethat · · focus · HN ↗
vintermann · · focus · HN ↗
leo4242 · · focus · HN ↗
You play White, and Jev plays Black. Rather than asking an LLM to generate move strings or JSON, the backend feeds all server-validated legal candidate moves into Vercel AI SDK's experimental_evaluate(). Jev picks Black's move and outputs its probability distribution across all legal candidates in a single forward pass (~300ms, ~$0.00004/move).
Github link: <a href="https://github.com/qibinlou/jev-chess" rel="nofollow">https://github.com/qibinlou/jev-chess
Give a try and let me know your thoughts! I am having lots of fun coming up with different chess strategies for Jev to try out.
haute_cuisine · · focus · HN ↗
ulcer · · focus · HN ↗
But it is essentially playing bullet chess right? No time to reason and is basically forced to go with it's knee jerk reaction. So maybe it plays reasonably against another average bullet chess player?
This is an instructive demo and makes me pause a bit when thinking about where I would trust using a classifier like this. Also there is no way to fine tune it right? So you are just left to its interpretation of the classification schema.
Building classifiers is hard and forces you into thinking about your problem space and your comfort level with type1 or type2 errors. I fear this will encourage sloppy work because all those decisions are black boxed.
Not like we were in a utopia careful ML applications before this.
wging · · focus · HN ↗
salicideblock · · focus · HN ↗
I definitely agree it's underexplained in type safe.ai's materials.
I have to assume it's a reference to the fast, heuristic, intuitive "system 1" process in humans, as opposed to the slow, procedural, reasoning "system 2".
This theory is recognized, among others, in Daniel Kahneman 2002 Nobel prize on Economics.
ostacke · · focus · HN ↗
ArafatMu · · focus · HN ↗